Evidence-to-action gaps expose tool-using agents to unsafe execution
SafeActBench reports that agents may act before evidence is complete despite high static task accuracy, especially in multi-step workflows. Evidence ledgers, provenance-bound tracking, trajectory checks, replayable traces, and approval gates provide concrete evaluation primitives, but harnesses must also test evidence manipulation, reward hacking, authorization abuse, publication changes, and adaptive agents.
Sources (2)
Updated Oct 10, 2026