Agent Benchmark Realism and Memory Challenges
AgentGym2 (ACL 2026) introduces de-idealized real-world agent benchmarks with noise, underspecification, and tool discovery, showing even GPT-5 and Gemini struggle. MemSearch-o1 addresses memory dilution in agentic search with token-level memory growth and path-based reasoning. These works highlight the gap between lab and real-world agent performance, driving need for robustness research.
Sources (2)
Updated Aug 10, 2026