Agentic AI security shifts toward access control, containment, and long-horizon behavior
Repository execution, credentials, mutable harnesses, multimodal prompt injection, controlled-decoding attacks, physical actions, and long-horizon collusion continue to evade simple refusal tests. Mechanistic-auditing results further show that interventions can damage correct baseline behavior, reinforcing the need for destination-resolved verification, scoped identities, least privilege, isolation, monitoring, and rollback; NVIDIA's platform approach still lacks broad independent evaluation.
Sources (5)
Updated Oct 5, 2026