AI Frontiers Digest

Anthropic Mythos Held Back as OpenAI GPT-5.5 Takes Agentic Lead; EdgeBench Reveals Agent Learning Scaling Law; LHTB Exposes Long-Horizon Limits

Anthropic Mythos Held Back as OpenAI GPT-5.5 Takes Agentic Lead; EdgeBench Reveals Agent Learning Scaling Law; LHTB Exposes Long-Horizon Limits

GPT-5.5 leads Terminal-Bench; new WildClawBench, FutureSim, CausalBench+, EvoPolicyGym, AgenticDataBench, CausalDS, UniClawBench, EvoCode expose evals gaps. Mythos withheld. ByteDance EdgeBench reveals log-sigmoid scaling law for agent learning. Long-Horizon Terminal-Bench (LHTB) shows even best model (Grok 4.5) only 15.2% pass@1, 29/46 tasks unsolved. Test-time compute budgets are a hidden confound (AISecurityInst). OpenAI released GPT-5.6 Sol with major vision gains (detection mAP 13.8→46.2, counting 73%) and GPT-5.6 Soul with strong persistence and token efficiency. New agent evaluation infrastructure AgentCompass and debugging tools for agent trajectories. RL for LLMs cycle continues; Microsoft scaling distribution-matching RL to large reasoning models. New: SearchOS-V1 tackles agent loop traps; SEED self-evolving on-policy distillation for agentic RL; Demystifying On-Policy Distillation formalizes pathologies; DeepLoop depth scaling for looped transformers. New agent eval critique shows harness evolution doesn't beat test-time scaling and generalizes poorly. GRASP introduces RL-based retrieval coordination for agentic RAG. ReOPD challenges 'more on-policy is better' dogma in distillation, offering 4× speedup. H^2SD hybrid self-distillation improves reasoning RL. GPT-5.6 Pro continues mathematical discoveries (Dinitz-Garg-Goemans conjecture). Self-evolving agents coevolve benchmarks with formal verifier grounding. New: LLMs Get Lost in Evolving User Intent formalizes blind spot in static evals. Experience Distillation retains 64.8% of ICL gains with 9.6x fewer samples. Tencent WorkBuddy Bench offers contamination-resistant coding-agent eval. AREX recursively self-improving agent achieves strong results on BrowseComp and HLE. New: Non-monotonic success-effort curve for Opus 5 on FrontierCode challenges monotonic improvement assumption. New: Progressive training pipeline from supervised to RL to agentic RL with implementation. New paper on overthinking in reasoning models. New: Molt PyTorch-native agentic RL training framework. New: JAXBench (Google) shows curated documentation lifts agent correctness from 5.8% to 37.3% for TPU kernel optimization. New: Grok Build introduces /deep-research command for multi-agent research. New: Offloading environment interactions to separate log file improves agentic generalization. New: Physics of multi-turn long-horizon planning paper dissects emergence across pre/post-training, finds suboptimal trajectories severely impair performance. New: EvoCode eval tests agents on evolving requirements; common failure is breaking existing functionality. New: Agent skills not always useful; regressions from skill description osmosis, grounding displacement, verification displacement. New: Kimi K3 weights released (2.8T MoE, 1M context, 2.5x intelligence-per-compute). New: Pass the Baton (Relay-OPD) tackles prefix failure in on-policy distillation, +5.73% over OPD with 50% less training. New: AI agents with $3,000 budget flunk open-ended AI research assignment—scored 2/6 and 1/6, failed to produce acceptable NeurIPS-level research. Key failures: poor judgment, inability to backtrack, impenetrable prose. Challenges RSI claims.

Sources (9)
Updated Aug 3, 2026
Anthropic Mythos Held Back as OpenAI GPT-5.5 Takes Agentic Lead; EdgeBench Reveals Agent Learning Scaling Law; LHTB Exposes Long-Horizon Limits - AI Frontiers Digest | NBot | nbot.ai