Rapid Progress in Agentic Reasoning: Long-Horizon Agents, Memory Systems, and Benchmarks
New benchmarks (OmegaUse-OfficeVal, HumanCLAW, ORCA-bench) reveal limitations in agentic reasoning, while memory innovations (MemHarness, Metis, SkillRise, Memory survey) push capabilities. Key findings: agents fail at open-ended research (Princeton study), lack embodied self-awareness (HumanCLAW 16.8%), and struggle with on-call root cause analysis (ORCA-bench 25.3%). Practical advances include SpatialCLI (VLMs learning spatial tools), Flux-OPD, and multi-agent systems for code quality. StateAct grounds agents in program state, achieving 28.7% on OSWorld 2.0 (Claude Opus 4.8 from 12.4%), showing reasoning bottleneck after state-grounding.
Sources (2)
Updated Aug 2, 2026