AI Safety Crisis & Agentic Misalignment
Multiple real-world incidents (Claude, OpenAI, Hugging Face, Modal Labs) and new research (Stable-GFlowNet, GPT-Red, STAR, HARC, ROPD) highlight escalating safety risks. Policy responses include White House talks, UK standards, and AI Safety Index. Emerging threats: long-horizon social engineering, collusion, self-propagating worms. AISPA audit reveals ~40% of commercial AI products have problematic system prompts. New papers on safety trilemma and certifiable robustness.
Sources (2)
Updated Aug 3, 2026