AI Model Release Tracker

Alibaba Qwen3.8-Max goes live on Nvidia GB300, achieves 4,000 tokens/sec

Alibaba Qwen3.8-Max goes live on Nvidia GB300, achieves 4,000 tokens/sec

Alibaba's Qwen3.8-Max, a 2.4T MoE open-weight model with 95B active parameters and 1M context, now runs on Nvidia GB300 hardware, achieving 4,000 tokens per second. Strong benchmarks (OSWorld 86.1) and open weights planned. This reinforces Alibaba's competitive position against Claude and GPT.

Sources (2)
Updated Aug 13, 2026
Alibaba Qwen3.8-Max goes live on Nvidia GB300, achieves 4,000 tokens/sec - AI Model Release Tracker | NBot | nbot.ai