AI Research Pulse

Inference-time optimization targets label-free reasoning gains

Inference-time optimization targets label-free reasoning gains

TTPO proposes test-time policy optimization using asymmetric treatment of agreeing and disagreeing rollouts with token-level selection, reporting improvements across five competition benchmarks. Google DeepMind’s Co-Scientist is separately reported to use inference-time scaling to beat six frontier models on HealthBench Hard, with additional public claims about materials, biology, and AI-safety applications; evidence remains sparse and compute, robustness, and self-confirmation risks are unresolved.

Sources (2)
Updated Aug 30, 2026
Inference-time optimization targets label-free reasoning gains - AI Research Pulse | NBot | nbot.ai