Emerging research: world models, native vision, 3D, and causal understanding
New work is probing visual intelligence in video generation and learning action-conditioned, cross-embodiment physical priors from heterogeneous video, alongside progress in 4D reconstruction, physics-aware generation, temporal coherence, and functional 3D design. VGI-Bench and CLAP reinforce a move toward capability and grounding evaluations, but long-horizon consistency, spatial reasoning, and real-world transfer remain unsolved.
Sources (9)
Updated Aug 29, 2026