Alibaba Qwen3.8-Max goes live on Nvidia GB300, achieves 4,000 tokens/sec
Alibaba's Qwen3.8-Max, a 2.4T MoE open-weight model with 95B active parameters and 1M context, now runs on Nvidia GB300 hardware, achieving 4,000 tokens per second. Strong benchmarks (OSWorld 86.1) and open weights planned. This reinforces Alibaba's competitive position against Claude and GPT.
Sources (2)
Updated Aug 13, 2026