NeuroByte Daily

Benchmarks for Long-Horizon Agent Controllers

Benchmarks for Long-Horizon Agent Controllers

Agent evaluation is expanding from pass/fail coding resolution to process, patch quality, tool use, trajectory, cost, and failure-phase metrics. LoopArena reports 24.69% strict success and 64.4% average inference-cost reduction; the 523-task study and CyberGym broaden coverage, but judge validity, methodology, and independent replication remain concerns.

Sources (4)
Updated Sep 1, 2026