AI Research Pulse

Post-training recipes, self-evolution, and prompt search are being challenged

Post-training recipes, self-evolution, and prompt search are being challenged

A preliminary Microsoft result argues that supervised fine-tuning may misalign models with subsequent reinforcement-learning objectives, while NPO reportedly reaches GEPA-like prompt-optimization performance with a single evolving lineage. J-Zero extends the trend to challenger–solver–judge co-evolution from zero data, and Rubric-to-Code explores localized credit assignment for interactive code generation. These findings could simplify or revise common optimization pipelines, but evidence remains preliminary and vulnerable to evaluator co-adaptation.

Sources (3)
Updated Aug 31, 2026
Post-training recipes, self-evolution, and prompt search are being challenged - AI Research Pulse | NBot | nbot.ai