Rubin platform expansion and CUDA readiness intensify inference competition
Vera Rubin and B300 anchor NVIDIA’s rack-scale strategy while CUDA 13.4 adds Rubin preview support, Windows on Arm64, MPS controls, Fabric Transport, and developer-tool improvements. NVLink Fusion and d-Matrix’s planned Raptor system extend NVIDIA’s platform to heterogeneous inference, while reported Pinterest B200/Dynamo deployment illustrates how KV-cache handling, disaggregated prefill/decode, routing, and orchestration increasingly determine economics.
Sources (21)
Updated Sep 11, 2026