Benchmarks for Long-Horizon Agent Controllers
Agent evaluation is expanding from pass/fail coding resolution to process, patch quality, tool use, trajectory, cost, and failure-phase metrics. LoopArena reports 24.69% strict success and 64.4% average inference-cost reduction; the 523-task study and CyberGym broaden coverage, but judge validity, methodology, and independent replication remain concerns.
Sources (4)
Updated Sep 1, 2026