AI Safety Vulnerabilities Escalate: Encrypted Reasoning Cracked, SkillJack, and Systemic Failures
A cross-session replay attack on encrypted reasoning blocks from Anthropic, OpenAI, and Google extracted 367 PII items and 182 credentials. SkillJack backdoor reduces safety detection from 98.5% to 11.4%. Other failures: Personalization Mirage (35-49% fabricated profiles), When Memory Lies (>2x death rate), and accountability gap (18% fewer errors caught when AI is anthropomorphized). These challenge assumptions about safety monitoring and privacy.
Sources (3)
Updated Aug 13, 2026