Agent harnesses and human collaboration become evaluation targets
Emerging research is testing how much agent behavior comes from the base model versus the surrounding harness, with traces from twelve public datasets reportedly compressible into a finite-state machine. Preliminary analysis of 7,400 trajectories suggests large context-length differences across harnesses, while ACE and persistent skill-library work emphasize the quality, structure, and reusability of agentic data.
Sources (1)
Updated Aug 29, 2026