Self-Improving Agent Systems Proliferate
Multiple new systems (AREX, Harness Handbook, OpenWorker, OpenForgeRL, Skill Self-Play) advance the paradigm of agents that improve their own code and planning. New Qwen-UI-Agent achieves SOTA on mobile benchmarks using online RL. ORCA-bench reveals frontier agents struggle with root cause analysis (25% on medium tasks), highlighting deployment challenges.
Sources (2)
Updated Aug 3, 2026