Agentic Inference Systems: CUDA Moat and TCO Trade-offs
AgentX InferenceXv3 analysis reveals agentic inference is a systems problem—prefix reuse, KV cache transfer, and routing dominate, not raw chip performance. DeepSeek V4 Pro results show TCO vs. interactivity trade-offs, with disaggregation sometimes hurting latency. Open-source vLLM/SGLang still trails vendor stacks like AMD's ATOM. Challenges the CUDA moat narrative and provides concrete benchmarks for agentic serving. Open Claude Code traces are a community resource. Updated with new developments: SpaceXAI adopts NVIDIA Vera CPU (1.8x task completion, 1.2TB/s bandwidth) for agentic workloads; NVIDIA BlueField-4 (800 Gb/s, 64 Arm cores) for AI factory networking; NVIDIA Groq 3 LPX for long-context interactivity on Vera Rubin. These reinforce that CPU orchestration and networking are now critical bottlenecks for agentic inference.