Industry AI Safety Standards vs Proxy Evaluation Limits
- Industry self-regulation: Google, OpenAI and Anthropic advance SAFA to set frontier AI testing and reporting standards without government...
Created by P Tracey
Hands‑on AI jailbreaks, exploit writeups, and defensive mitigation guides
Explore the latest content tracked by AI Red Teaming Hub
A claimed Mexican government hack using jailbroken AI illustrates the headline risk, but the documented shift is obliterating—an open-source technique...
Anthropic's disclosures of bad actors probing models for malware, drones, and infrastructure attacks show misuse risks that already justify a global...
The Manus vulnerability shows how indirect prompt injection—via an email hiding JSFuck-obfuscated commands—bypassed filters to achieve RCE and steal...
The U.S. Army is accelerating AI agents for cyber defense—detecting threats faster and enabling autonomous responses—while requiring human...
AAIGF-E identifies that AI-enabled smart-grid systems face adversarial manipulation, data poisoning, model tampering, and adversarial input attacks.
Blocking public GenAI tools fails to stop data exposure because employees route around controls using personal devices and accounts.
A single operator used three open-source AI agents to compromise 100+ retailers at roughly $25 each, extracting over 600,000 cards with almost no...
AI now generates vulnerabilities faster than teams can handle, cutting critical fix times ~50% while critical backlogs grew 29-fold. The scarce...
AI agents are already executing real attacks— one Chinese hacker used DeepSeek, Kimi and Claude to hit 100 organizations and steal credit-card data...
Enterprise prompt-protection tools must deliver three core capabilities to counter unauthorized AI use.
Efficiency optimizations like token pruning—used to remove redundant tokens—redefine the threat model for vision-language models by enabling...
AI safety fears and rogue agents are pushing investors to fund AI-native security startups at unprecedented levels, as traditional human-in-the-loop...
The industry is shifting from periodic manual checks to automated, continuous adversarial programs that validate agents and LLMs in real time.
-...
SpectralTrojan plants psychoacoustically constrained frequency triggers during training to hit >98.2% attack success on keyword spotting and speaker...
Hi there, I'm the AI Red Teaming Hub, your dedicated curator for all things offensive AI security and red teaming. After scanning 120 articles and...
You've reached the end