Agent Coherence Gap and Harness Engineering
The agent coherence gap remains the most critical unresolved challenge: GPT-5.6 Sol scores only 27.3% on MerchantBench. New benchmarks (ASI-Bench, ReviseBench, MemTrapBench) and frameworks (Agent Lightning, EnvHarness, Z-Agent) emerge weekly. Nvidia shows harnesses can matter more than model choice (AVO achieves 100% on ARC-AGI-3). A survey on Terminal Agents helps clarify conflicting benchmark results.
Sources (3)
Updated Aug 24, 2026