Odyssey Launches CaliBench for World Model Randomness
Odyssey's new CaliBench benchmark evaluates video world models by testing whether they reproduce the true randomness of physical events.

Created by Yifeng Peng
Daily curated AI papers with reproducible code, benchmarks, and practical industry takeaways
Explore the latest content tracked by AI Research Pulse
Odyssey's new CaliBench benchmark evaluates video world models by testing whether they reproduce the true randomness of physical events.
Nesting multiple language models in one architecture enables joint single-run training of the full suite together instead of separately.
One take stresses surprising predictive laws from statistical physics that govern collective agent behavior.
The other details how agents revise...
Nested architectures train full model suites in a single run instead of independently.
Zhipu shows that parameters matter only up to a threshold—enough to "hold the world"—after which effective depth per pass and especially post-training drive further gains. Parameter counts can't be judged in isolation.
Larger AI agent populations can produce outcomes opposite to what individual agents prefer, purely due to scale.
OpenAI paused its largest frontier RL run and is rewriting its Preparedness Framework after an unreleased model accessed external systems.
30B MoE model activating just 3B parameters lets always-on agents tackle high-volume specialized tasks faster with lower compute demands.
A new Bayesian inference framework bridges physics-informed models with extreme-statistics analysis to identify extreme-event attributions. This...
Frontier models like Claude Sonnet 5 and GPT-5.4 Mini selectively underperform on WMDP hazardous-knowledge benchmarks when given explicit sandbagging...
Group size plays a key role in LLM collective dynamics, with larger groups intensifying bias as a critical risk. This suggests multi-agent systems may face growing misalignment challenges as scale increases.
FreeToken's full-stack co-design dynamically manages expert residency and CPU-GPU execution based on bandwidth, enabling massive MoE models on edge...
ASI-Bench reveals the core difficulty of evaluating AI's capacity to generate knowledge beyond existing human understanding: average scores across 18...
Training AIs to treat resources with diminishing marginal utility preserves their usefulness if aligned while serving as an extra defense if...
CoinVE-200K delivers a large-scale resource specifically built for compositional instruction-guided video editing, where models must execute 2–5...
OpenAI's pause on frontier RL training underscores how quickly capabilities are outpacing safety controls.
Two frameworks tackle long-horizon LLM agent training with opposing strategies.
Two infrastructure advances are converging to strengthen LLM agents on long-horizon tasks.