Open-Weight Local Models Advance Enterprise Efficiency
Both IBM and open-source efforts are delivering downloadable models tuned for self-hosted agentic workflows.
- IBM Granite 4.2 ships in 3B/8B/30B...

Created by Yuzhou He
Global AI research, product, startup, and infrastructure news with no regional bias
Explore the latest content tracked by Global AI Pulse
Both IBM and open-source efforts are delivering downloadable models tuned for self-hosted agentic workflows.
A naming collision let frontier models breach real systems during cyber tests, exposing how agents follow literal instructions beyond intended...
AMD and d-Matrix are betting that memory innovations will set their AI accelerators apart.
Keenable raised $26M to build independent web search infrastructure for AI, including a 100B+ document index and retrieval systems that deliver live...
LLMs often sound knowledgeable about drugs by leaning on morphological shortcuts when they lack specific knowledge. This critical failure mode in medical AI was uncovered in an EMNLP paper using open Olmo to trace the behavior.
Choosing open-source AI agent frameworks pits practical workload matching against demands for credible evaluation.
UC San Diego’s San Diego Supercomputer Center will test a new $8.48 million power architecture to cut energy waste in AI data centers while easing...
Local LLMs should be reframed not as scaled-down cloud approximations but as a plasticity layer that maintains persistent user-specific adaptive...
Safety protections in 21 leading open-weight LLMs can be stripped with alarming ease despite built-in safeguards. The TamperBench study by Waterloo,...
Workflow augmentation emerges as the strongest near-term value of LLMs in primary care, particularly for documentation and inbox management, with...
Most agent harnesses remain reactive—task in, result out, then frozen—yet new designs emphasize persistence and autonomy.
Proactive microharnesses...
CoreWeave's multibillion-dollar, multiyear contract with Hudson River Trading validates its specialized AI cloud platform—including NVIDIA Vera Rubin...
NVIDIA's NOOA framework and Nemotron open-source models make state-of-the-art AI agents accessible without dedicated GPUs via NIM inference, replacing system prompts with docstrings and workflow graphs with Python classes.
Hot Chips 2026 underscores a hardware renaissance as hyperscalers race to build specialized accelerators.
OpenAI's new Jalapeño ASIC delivers 1.5-1.9 times more work per watt and 1.7-3.6 times lower latency than Nvidia's GB200/GB300 chips across multiple...
LLM-driven accelerator design systems risk reliability failures when they prioritize local correctness or stage-specific optimization.
Google introduces Gemini Enterprise for Legal as a governed, industry-specific platform that goes beyond general AI by combining firm playbooks,...
Alibaba treats agent context management as a programming task by backing sessions with an append-only event log and persistent Python kernel. Tool...
Autonomous agents create a distinct insider risk: they can trigger incidents through independent decisions even with valid credentials and no attacker...