Xiaomi XR-1: Human Video Pretraining Challenges Teleoperation Scaling Dead End
Key Questions
What is Xiaomi's XR-1 model and how does it pretrain?
Xiaomi's XR-1 model pretrains on 100K hours of human video data using UMI without any robot data, then fine-tunes with 7.2K hours of real robot data. This approach achieves 75-85% success rates compared to π0.5's 40-53%.
How does XR-1 challenge existing views on teleoperation data scaling?
XR-1 demonstrates a clean scaling law and zero-shot transfer potential, directly challenging the idea that scaling teleoperation data is a dead end. It offers a concrete path to breaking the data bottleneck in embodied AI.
What performance does XR-1 achieve after fine-tuning?
After fine-tuning, XR-1 reaches 75-85% success rates on tasks, significantly outperforming π0.5's 40-53%. The model uses minimal robot data following extensive human video pretraining.
Xiaomi's XR-1 model pretrains on 100K hours of human video (UMI) without any robot data, then fine-tunes with 7.2K hours of real robot data, achieving 75-85% success vs π0.5's 40-53%. This clean scaling law and zero-shot transfer potential directly challenges the argument that scaling teleoperation data is a dead end, offering a concrete path to breaking the data bottleneck in embodied AI.