NeoHorse-1 — harness-mediated agentic post-training and self-improvement
NeoHorse-1 combines heterogeneous routing, validated interaction traces, evaluation-driven selection, and on-policy distillation to improve 4B and 9B models. Its strongest contribution is an apparently practical training harness, while claims of recursive self-improvement remain preliminary pending independent reproduction, cross-generation gains, and checks for bias accumulation or capability regression.
Sources (2)
Updated Sep 9, 2026