AI Research Pulse

Coding-agent benchmarks confront real-world reliability

Coding-agent benchmarks confront real-world reliability

Coding agents reportedly surpass 95% on SWE-bench, but AgentLens finds 10.7% lucky passes, while unreleased task variants can cause large pass@3 drops and ranking changes. Regressions, literal prompt following, poor repository hygiene, contamination, and review overhead are shifting attention from leaderboard scores to dependable software-workflow integration.

Sources (2)
Updated Aug 29, 2026
Coding-agent benchmarks confront real-world reliability - AI Research Pulse | NBot | nbot.ai