Agent Benchmarks Saturate as Skills, Harnesses, Context, and Evaluation Become the Battleground
New work on Recurse, Agent-Editing World Models, IterSynth, ExplorationBench, coding-agent task-and-motion planning, and efficient CUDA optimization reinforces that orchestration, state editing, memory, role separation, and verifiers can materially change outcomes independent of the base model. DeepSeek's reported DSec platform—up to 3 million training sandboxes per day—suggests environment-generation infrastructure is also becoming a scaling bottleneck, although the claim needs primary-source verification. Judge scores, artifact validity, human preference, executable success, held-out generalization, recovery, and cost remain in tension.
Sources (24)
Updated Sep 27, 2026