Long-horizon multimodal mem/world models
Key Questions
What performance does Flash-WAM deliver?
Flash-WAM achieves a 23x speedup in long-horizon multimodal memory and world-model workloads.
How does ARROW reduce forgetting compared with DreamerV3?
ARROW cuts forgetting by 20x relative to DreamerV3 while maintaining long-horizon performance.
What TSR score is reported for SpatialWorld GPT-5?
SpatialWorld GPT-5 records 17.4% TSR on the evaluated spatial-reasoning tasks.
What key finding does the Imagined Rollouts paper present?
It diagnoses that world-model rollouts are primarily kinematic rather than dynamic, with DreamerV3 rollouts approximately 100x more kinematic than real dynamics.
What efficiency gains does Light-Omni provide for video agents?
Light-Omni delivers a 12.1x speedup and 2.6x memory reduction by favoring reflex over reasoning with long-term memory.
What tasks does the RoboDojo benchmark cover?
RoboDojo supplies a unified sim-and-real benchmark containing 42 simulation and 18 real-world robot-manipulation tasks.
What is MIRA and its runtime performance?
MIRA is a multiplayer interactive world model trained on Rocket League that runs at 20 fps on a single GPU.
How does SIEVE affect data needs for VLA imitation learning?
SIEVE's structure-aware selection reduces required training data by 50% while preserving policy performance.
Multimodal advances: EVA-Client, Light-Omni (12.1x speedup), AlayaWorld, Flex-Forcing, SIEVE (50% data reduction), RynnWorld-Teleop, Nemotron-Labs-Diffusion, Signed MaxSim. Also MIRA, LaMem-VLA, LingBot-Video, RoboDojo, OmniTacTune. No new updates today.