Core ML Research

Self-Improving AI with Harness & Weight Updates

Self-Improving AI with Harness & Weight Updates

Key Questions

What is SIA and what performance gains does it report?

SIA unifies harness and weight updates for self-improving agents, achieving a 56.6% gain on LawBench and 502% denoising improvement. It forms part of a growing thread of papers on self-improving AI systems.

How does Bayesian-Agent extend self-improving agents?

Bayesian-Agent builds on SIA by incorporating Bayesian posterior updates, improving performance from 80% to 95% on SOP-Bench.

What gains do Role-Agent and EEVEE provide?

Role-Agent delivers a 4% gain while EEVEE achieves 10-24 point improvements through router-prompt co-evolution in self-improving setups.

What is HarnessBridge and where does it perform well?

HarnessBridge introduces a learnable bidirectional controller and shows strong results on Terminal-Bench and SWE-bench.

What benchmarks address agent memory issues?

MemSyco-Bench evaluates sycophancy in agent memory, while A-TMA tackles ghost memory using state-aware overlays and evidence packets for a 0.240 gain on LTP.

What new environments test long-horizon agents?

AgenticSTS provides a bounded-memory testbed using Slay the Spire 2, and EvoPolicyGym benchmarks autonomous policy evolution where GPT-5.5 leads.

Which tools support agent state management?

Stanford researchers released a checkpointing tool for files, DB, and KV cache modeled after Git, alongside InternVideo3 for agentifying foundation models.

What other self-improvement methods are mentioned?

Additional work includes agent memory as data management, Cursor's 1.5T model training on Colossus, agentic evolution for hardware-aware compression, AutoTrainess for automated LM post-training, and DiscoPER for autonomous scientific discovery.

A growing thread of papers on self-improving agents. SIA unifies harness and weight updates (56.6% gain on LawBench, 502% denoising). Bayesian-Agent extends with Bayesian posterior (80→95% on SOP-Bench). New additions: Role-Agent (4% gain), Self-Harness, EEVEE (10-24 point gains via router-prompt co-evolution). InternVideo3 agentifies foundation models. Latest: HarnessBridge (learnable bidirectional controller, solid on Terminal-Bench and SWE-bench). Also: agent memory as data management, Cursor training 1.5T model on Colossus, agentic evolution for hardware-aware compression, AutoTrainess (automates LM post-training), MemSyco-Bench (sycophancy in agent memory), DiscoPER (autonomous scientific discovery). New: EvoPolicyGym benchmarks autonomous policy evolution in interactive environments (GPT-5.5 tops). New: AgenticSTS provides bounded-memory testbed for long-horizon agents using Slay the Spire 2, showing gains from strategic skills. Also noted: a new tool for agent state management (checkpointing files, DB, KV cache like Git) from Stanford researchers. New: A-TMA addresses ghost memory with state-aware overlay and evidence packets (0.240 absolute gain on LTP).

Sources (9)
Updated Jul 8, 2026