AI Theory Frontier

Label-Free Test-Time Optimization and Distillation

Label-Free Test-Time Optimization and Distillation

TTPO reportedly matches supervised OPSD while improving Qwen3-1.7B on competition mathematics through filtered pseudo-labels and asymmetric agreement handling. Self-OPD, rubric-to-code credit assignment, and DART-SD extend self-distillation and localized credit assignment to flow matching, coding, and multi-turn tool use, but cross-domain generalization remains unresolved.

Sources (2)
Updated Sep 1, 2026
Label-Free Test-Time Optimization and Distillation - AI Theory Frontier | NBot | nbot.ai