Agent Harnesses, Skills, Evaluation & the AI Security Supply Chain
Persistent coding environments, self-hosted sandboxes, reusable skills, tool-use research, compiler testing, documentation and scientific-diagram evaluations, ProVer credit assignment, and AI-generated security reports make permissions, trajectory inspection, recovery, triage, and artifact verification central concerns. Skill swarms and semantic code-search discussions add practical signals, but judge dependence, benchmark overfitting, adversarial robustness, calibration, and real-world transfer require stronger evidence.
Sources (49)
Updated Oct 6, 2026