AI Breakthrough Digest

Agent Evaluation and Autonomy: ASI-Bench and Beyond Final Scores

Agent Evaluation and Autonomy: ASI-Bench and Beyond Final Scores

A new wave of benchmarks and training methods is redefining agent capabilities. 'Beyond Final Scores' finds agents are 'engineering optimizers' not 'autonomous researchers'. ASI-Bench (60 tasks, 11 domains) reveals a sharp performance drop when human guidance is removed, confirming heavy scaffolding. Agent Lightning v1.0 achieves 14.6-point gain on SWE-bench with harnessed agentic RL using only 6K examples. Agentic ESOpt uses evolution strategies for fine-tuning long-horizon agents, achieving 6.69% improvement on WebArena-Lite with minimal GPU memory. These developments challenge hype and provide concrete directions for improvement.

Sources (3)
Updated Aug 20, 2026