Applied AI Brief

AI Compute Race: Custom Chips, Efficient Inference, and Physical AI Capital

AI Compute Race: Custom Chips, Efficient Inference, and Physical AI Capital

Key Questions

Why are AI labs developing custom chips?

DeepSeek is developing its own AI chip to reduce Nvidia dependency, similar to Anthropic's Samsung partnership. This vertical integration trend challenges Nvidia's dominance and impacts supply chains.

What efficiency targets is Google pursuing with its Frozen v2 chip?

Google's Frozen v2 aims for 6-10x token efficiency over TPUs by embedding Gemini architecture in silicon, targeting 2028 deployment despite flexibility trade-offs.

What acquisitions or investments signal AI chip verticalization?

Apple is exploring AI chip startup acquisitions for inference and Private Cloud Compute. Siemens acquired Precision Innovations for AI-powered chip planning to reduce design iterations.

How are companies addressing AI inference hardware specialization?

Etched AI chip startup reached $10.3B valuation focusing on MoE/Mamba architectures. Huawei shipped Attention-FFN Disaggregation for optimized MoE serving latency and cost.

What infrastructure expansions support the AI chip race?

Micron begins $9.3B Japan plant expansion. AMD announced ROCm.AI claiming 3.3x inference improvement, while Google hoards TPUs for AGI research amid supply constraints.

The efficiency race spans architecture, training, memory, interconnects, cloud pricing, workload-specific silicon, and physical AI. SMELT reports 6.8–18.0% training-FLOP savings and HeadWiseKV targets KV-cache residency; a reported $12.9B Nvidia-Hugging Face acquisition would potentially reshape the open-model and compute ecosystem but remains unconfirmed. Deployed-workload evidence and neutrality implications remain key questions.

Sources (17)
Updated Sep 4, 2026
Why are AI labs developing custom chips? - Applied AI Brief | NBot | nbot.ai