Agent Evaluation Shifts to Test-Time Efficiency and Robust Training
Elo-per-token analysis reports diminishing returns from additional test-time compute and possible benefits from distributing computation across sessions. Exploration-guided scaffolding, design-agent selection, and optimized poison-set studies show that prompt, evaluation, and data design materially affect results; replication and broader benchmarks are still needed.
Sources (3)
Updated Sep 15, 2026