AI Breakthrough Tracker

Robotics and Vision-Language Model Advances

Robotics and Vision-Language Model Advances

Key Questions

What advances have been made in vision-language-action (VLA) models?

VLA models have surged with NVIDIA MotionBricks offering over 350k skills at 15k FPS and 2ms latency. Qwen-VLA reached 97.9% on LIBERO, and PRTS VLA achieved 95.9% in real-world tests.

What is the HIW-500 dataset and its significance?

HIW-500 is an open-source humanoid dataset with over 500 hours of data from 12 homes. It supports training and evaluation of robotic systems in realistic household environments.

How does In-Context World Modeling benefit robotic control?

In-Context World Modeling allows VLA adaptation without fine-tuning by leveraging existing trajectories. It also enables predictable prevention of hallucinations in world models using as few as 50 trajectories.

What is PhysiFormer designed to do?

PhysiFormer uses a diffusion transformer on 3D meshes to simulate mechanics directly in world space. It advances physical understanding for robotics and vision-language applications.

What funding has AMI Labs received for world model research?

AMI Labs secured a $1 billion seed round to advance JEPA-based world models. Yann LeCun has publicly advocated for this direction in AI development.

VLA models surge. NVIDIA MotionBricks (350k+ skills, 15k FPS, 2ms latency). Qwen-VLA (97.9% LIBERO), PRTS VLA (95.9% real-world), ACE-Ego-0 unifying egocentric human and robotic data. HIW-500 open-source humanoid dataset (500+ hours, 12 homes). Qwen-AgentWorld language world model (397B beats GPT-5.4 on environment prediction). In-Context World Modeling enables VLA adaptation without fine-tuning. Hallucination in world models predictable and preventable with as few as 50 trajectories. PhysiFormer simulates mechanics in world space via diffusion transformer on 3D meshes. Yann LeCun argues for world models; AMI Labs $1B seed for JEPA. LTX-2.3 targets VFX with controllable video gen at 1/5-1/10 cost.

Sources (3)
Updated Sep 23, 2026
What advances have been made in vision-language-action (VLA) models? - AI Breakthrough Tracker | NBot | nbot.ai