Agentic AI & Simulation

Self-Evolution & Runtime: Continual Improvement and Experience Distillation

Self-Evolution & Runtime: Continual Improvement and Experience Distillation

New: Simulator collapse in multi-agent RL — single LLM simulator leads to mode collapse and overfitting; Verbalized Sampling and Co-Training mitigate. Latent On-Policy Self-Distillation (LOPD) makes teacher's context learnable end-to-end, outperforming RLVR/OPSD/Skill-SD with <30% rollout budget. Other signals: SEARL, GECCO paper on evolving multi-agent systems, recursive self-improvement, experience distillation.

Sources (2)
Updated Aug 18, 2026