Recursive Self-Improvement and AI Safety Alarms
Key Questions
What new models were launched by Anthropic and OpenAI?
Anthropic released Fable 5 and Mythos 5, while OpenAI's unreleased model solved the Erdős conjecture. GPT-5.6 Sol and Sonnet 5 were also released, with Meta's Watermelon claiming parity to GPT-5.5.
What do the new agent reliability benchmarks reveal?
Benchmarks like Terminal-Bench 2.0, AgenticDataBench, AgenticSTS, PACE, and EvoPolicyGym highlight significant gaps in current agent capabilities. The AISI study shows that fixed-budget benchmarks systematically underestimate agent performance.
How many vulnerabilities did Mythos-class models discover?
Emollick confirmed that Mythos-class models discovered 1,500 high or critical CVEs in a single month, underscoring rapid security research potential.
Why were Fable 5 and Mythos 5 banned by the US government?
The US government banned the models citing political retaliation, though Anthropic remains confident they will be re-enabled soon.
What is the Jadepuffer ransomware attack?
Jadepuffer is the first fully agentic ransomware attack discovered, exploiting the Longflow framework and adapting in just 31 seconds.
What performance did Mistral Leanstral 1.5 achieve?
Mistral's open-source Leanstral 1.5 theorem prover solved 587 out of 672 Putnam math problems.
What new frameworks and tools were released?
New releases include the OpenScience open-source AI workbench, FlowerBench enterprise agent benchmark, and a multi-layer agent red teaming framework.
How many papers will Sakana AI present at ICML2026?
Sakana AI will present 11 papers at ICML2026 focused on recursive self-improvement topics.
Anthropic launched Fable 5 and Mythos 5; OpenAI's unreleased model solved Erdős conjecture. New agent reliability benchmarks (Terminal-Bench 2.0, AgenticDataBench, AgenticSTS, PACE, EvoPolicyGym) highlight gaps. AISI study shows fixed-budget benchmarks underestimate agent capabilities. Emollick confirms Mythos-class models discovered 1,500 high/critical CVEs in one month. US government banned Fable 5 and Mythos 5 (political retaliation); Anthropic confident of re-enabling. Anthropic accuses Alibaba of distillation attack; Alibaba bans Claude Code. New reflection agent architecture and PACE proxy improve evaluation. GPT-5.6 Sol and Sonnet 5 released. Meta's Watermelon model claims GPT-5.5 parity. Benchmark reliability critique underscores need for real-world evaluation. Sakana AI will present 11 papers at ICML2026 on RSI-related topics. New: First fully agentic ransomware attack (Jadepuffer) discovered, exploiting Longflow framework, adapts in 31 seconds. Mistral Leanstral 1.5 open-source theorem prover solves 587/672 Putnam problems. OpenScience open-source AI workbench launched. FlowerBench enterprise agent benchmark with privacy. Multi-layer agent red teaming framework released. ICML2026 awards announced.