Reality checks: RoboDojo, ALLEX drill test, and H2R-Bench expose foundation model and world model limitations
RoboDojo benchmark shows best generalist policy at 8.80% success vs 76% human teleoperation. ALLEX drill test shows 100% grasp but only 50% alignment. H2R-Bench finds even top world models fail at embodiment consistency and functional interaction. All highlight precision bottlenecks and the gap between foundation models and real-world manipulation.
Sources (4)
Updated Aug 17, 2026