Agentic Reasoning and Self-Evolution
Key Questions
What progress is occurring in agentic reasoning benchmarks?
Rapid advances include papers on long-horizon agents, self-evolving benchmarks with Lean verifiers, and RL for scientific discovery via DiscoBench.
How are self-improving agent loops being engineered?
Practical guides cover agent memory architecture and loop engineering, with new work on self-improving agents using coevolving benchmarks and environment-free synthetic data.
What does the DocOps benchmark reveal about document agents?
DocOps exposes critical failures such as long-term state tracking collapse, shallow verification, and destructive metadata editing in autonomous document workflows.
What new resources support video memory and RL for novelty?
ReflectWorld-MM addresses video memory, while Roberta Rail's talk focuses on RL for novelty in agentic systems.
Which paper enables long-horizon reasoning through memory?
Programmatic Memory Enables Long-Horizon Reasoning (PRO-LONG) shows how programmatic memory supports sustained perception and reasoning in extended tasks.
Rapid progress in agentic reasoning benchmarks and self-improving loops. New papers on long-horizon agents, self-evolving benchmarks with Lean verifiers, and RL for scientific discovery (DiscoBench). Practical guides on agent memory architecture and loop engineering. New today: @omarsar0's self-improving agents with coevolving benchmarks, Roberta Rail's talk on RL for novelty, ReflectWorld-MM for video memory, environment-free synthetic data for API agents. DocOps benchmark exposes critical failure modes in document agents (long-term state tracking collapse, shallow verification, destructive metadata editing) — directly relevant for enterprise document workflows.