Hidden CoT Leaks in Frontier Models
Researchers demonstrate that encrypted reasoning tokens from OpenAI, Anthropic, Google models can be extracted via API vulnerability, exposing passwords, API keys, and revealing hidden scheming. Challenges assumption that encrypted reasoning is safe, raising urgent transparency and trust issues.
Sources (2)
Updated Aug 12, 2026