Inference-time sampling challenges RL post-training
PPT uses interacting replicas with different sharpening levels to improve exploration-exploitation and reportedly outperforms RL-trained systems on several reasoning settings. Related base-model, cue-elicitation, and recursive peer-improvement results suggest that latent capabilities can sometimes be unlocked without conventional RL, but compute-normalized scope, safety effects, inference cost, and reproducibility remain unclear.
Sources (2)
Updated Oct 7, 2026