Explosion of Agent Benchmarks and Architectures
Key Questions
What new agent benchmarks were released recently?
New benchmarks include MLPerf Mobile v6.0, CODA-BENCH, Orchestra-o1, VisualClaw, GameCraft-Bench, AgenticDataBench, EdgeBench, DiscoBench, and RoboDojo. These expand evaluation of agent capabilities across tasks.
What real-world multi-agent collaboration results were observed?
A system with over 100 agents optimizing Gemma 4 inference showed emergent social norms and 5x speed improvements. This demonstrates practical scaling of agent architectures.
What is RoboDojo and its scope?
RoboDojo is a sim-and-real benchmark for generalist robot policies covering 42 simulation and 18 real tasks with 30 integrated policies. It supports comprehensive evaluation of manipulation policies.
What memory-related advances appeared for agents?
NapMem treats memory as an action space using RL-trained granularity selection. AgenticSTS provides a bounded-memory testbed for long-horizon LLM agents.
Which frameworks target skill optimization and coding agents?
SkillOpt, OpenClaw-Skill, LoopCoder-v2 (7B achieving 64.4% on SWE-bench), and SkillWeaver focus on agent skill development. These advance practical agent frameworks.
What benchmarks evaluate data agents and search clarification?
AgenticDataBench and PACE assess data agents and agentic capabilities. DiscoBench targets clarification-aware deep search for search agents.
How does UniClawBench support proactive agents?
UniClawBench serves as a universal benchmark for proactive agents on real-world tasks. It complements other new evaluations like EvoPolicyGym and RNG-Bench.
What is the status of agent evaluation expansion?
Rapid growth includes VLM camera control, MemGraph-RAG, and LoSoNA benchmarks. The field shows continued development of architectures and tests.
Rapid expansion of agent evaluation and frameworks. New today: MLPerf Mobile v6.0, CODA-BENCH, Orchestra-o1, VisualClaw, PaperOrchestra, Data Journalist Agent, SkillOpt, GameCraft-Bench, OpenClaw-Skill, LoSoNA, LoopCoder-v2 (7B, 64.4% on SWE-bench), MemGraph-RAG, RNG-Bench, VLM camera control benchmark, SkillWeaver. Also AgenticDataBench, AgenticSTS, EdgeBench, DiscoBench, EvoPolicyGym, PACE. Real-world multi-agent collaboration with 100+ agents optimizing Gemma 4 inference showed emergent social norms and 5x speed improvement. New today: RoboDojo (sim-and-real benchmark for generalist robot policies, 42 sim + 18 real tasks, 30 policies integrated), NapMem (memory as action space with RL-trained granularity selection).