Agent Evaluation and Benchmarking Evolution
New benchmarks (SWE-Bench ProMax, Evo-Bench, SWE Odyssey) and evaluation frameworks (Beyond Final Scores) reveal that agents act as engineering optimizers rather than autonomous researchers, with process bottlenecks and misleading experience reuse. This shifts focus from final scores to nuanced agent behavior in long-horizon tasks.
Sources (5)
Updated Aug 18, 2026