Zero-RL Scaling to Trillion Parameters and Latent Reasoning
Key Questions
What scale did Ring-Zero achieve with zero RL?
Ring-Zero scaled zero reinforcement learning to 1 trillion parameters at Ant Group. It produced five emergent behaviors including structured reasoning and context anxiety.
What does context anxiety refer to in Ring-Zero?
Context anxiety describes the model's self-managing of compute budget during reasoning. It emerged as one of five key behaviors from the scaling experiment.
What does SLPO enable for latent reasoners?
SLPO brings outcome-reward RL to latent reasoners. It solves per-step likelihood estimation and adaptive stopping without explicit CoT token costs.
Ring-Zero reports trillion-parameter zero-RL scaling with emergent structured reasoning, self-verification, context anxiety, and distinct discovery/sharpening phases. ProVer's segment-level credit assignment, SLPO for latent reasoners, and new density-aware reward aggregation extend the agenda toward more efficient post-training, though judge reliability and scaling costs remain open.