AI Frontiers Digest

Zero RL Scaling to Trillion Parameters: Emergent Reasoning and Context Anxiety

Zero RL Scaling to Trillion Parameters: Emergent Reasoning and Context Anxiety

Key Questions

What scale did Ring-Zero achieve with zero RL?

Ring-Zero scaled zero reinforcement learning to 1 trillion parameters at Ant Group. It produced five emergent behaviors including structured reasoning and context anxiety.

What does context anxiety refer to in Ring-Zero?

Context anxiety describes the model's self-managing of compute budget during reasoning. It emerged as one of five key behaviors from the scaling experiment.

What does SLPO enable for latent reasoners?

SLPO brings outcome-reward RL to latent reasoners. It solves per-step likelihood estimation and adaptive stopping without explicit CoT token costs.

Ring-Zero (Ant Group) scales zero reinforcement learning to 1 trillion parameters, yielding five emergent behaviors: structured reasoning, self-verification, context anxiety (self-managing compute budget), and two-phase training dynamics (discovery then sharpening). Validates the 'bitter lesson' and challenges assumptions about human engineering. A major milestone in scaling RL for LLMs. New: Paper on understanding reasoning from pretraining to post-training finds joint scaling law and compute allocation insights, complementing Ring-Zero. New: SLPO brings outcome-reward RL to latent reasoners, solving per-step likelihood and adaptive stopping—potential paradigm shift for efficient reasoning without explicit CoT token cost.

Sources (2)
Updated Jul 23, 2026
What scale did Ring-Zero achieve with zero RL? - AI Frontiers Digest | NBot | nbot.ai