Schema Harness Saturates ARC-AGI-3; Harness Evolution Debate Intensifies
Key Questions
What does the Schema harness achieve on ARC-AGI-3?
The custom Schema harness reaches saturation on the challenging ARC-AGI-3 benchmark, providing strong evidence for the value of agent scaffolding and test-time compute.
What does the paper 'Rethinking the Evaluation of Harness Evolution for Agents' conclude?
It finds that automatic harness evolution does not consistently outperform simple test-time scaling. The work argues researchers may be optimizing the harness rather than the underlying model.
What efficiency improvement comes from Recursive Harness Self-Improvement?
The algorithm enables continual learning in model-harness co-evolution and reduces costs by 60%. It offers a practical approach to ongoing harness refinement.
A custom harness called Schema achieves saturation on the notoriously hard ARC-AGI-3 benchmark. Significant signal for agent scaffolding and test-time compute. New paper 'Rethinking the Evaluation of Harness Evolution for Agents' challenges automatic harness evolution, finding it doesn't consistently beat simple test-time scaling. Recent tweet by @CharlesVardeman reinforces the argument that we are optimizing the wrong thing—the harness, not the model. New: Recursive Harness Self-Improvement paper offers a practical algorithm for continual learning in model-harness co-evolution, achieving 60% cost reduction.