AI Breakthrough Radar

Frontier Model Capabilities and Benchmarks

Frontier Model Capabilities and Benchmarks

Key Questions

What mathematical achievement did GPT-5.6 demonstrate?

GPT-5.6 used a 10-page prompt to close a 30-year gap in convex optimization, producing novel mathematical results. This challenges views that LLMs cannot generate original research.

How does Kimi K3 compare to leading models?

Kimi K3 (2.8T parameters) beats GPT-5.6 Sol and Opus 4.8, signaling a potential pricing moat collapse. It highlights rapid capability gains from Chinese labs.

What new large models are entering the frontier race?

Alibaba previewed Qwen3.8-Max, a 2.4-trillion-parameter multimodal model. Elon Musk announced Grok 4.6 (2T parameters) will finish training next week.

How does Muse Spark 1.1 perform against GPT-5.6?

Muse Spark 1.1 beats GPT-5.6 on SciCode with 58% accuracy and nearly matches Fable 5 and Gemini 3.1 Pro. It also shows strong creative outputs in video generation.

What does the competitive landscape indicate for 2026?

Rapid releases of massive open-weight models like Qwen3.8 and Kimi K3 are shifting the landscape. Pricing pressure and capability parity are accelerating across labs.

GPT-5.6 used a 10-page prompt to close a 30-year gap in convex optimization, demonstrating LLMs as research tools capable of novel mathematical results. This milestone challenges the notion that LLMs cannot produce original research. Also, Muse Spark 1.1 beats GPT-5.6 on SciCode (58%) and nearly matches Fable 5 and Gemini 3.1 Pro. Kimi K3 (2.8T params) beats GPT-5.6 Sol and Opus 4.8, signaling pricing moat collapse. These developments highlight rapid capability advances and shifting competitive landscape.

Sources (6)
Updated Jul 20, 2026