Embodied AI Digest

Mimic's FLUX-mimic Achieves 95% Success on Soft-Body Kitting, Challenges VLA Paradigm

Mimic's FLUX-mimic Achieves 95% Success on Soft-Body Kitting, Challenges VLA Paradigm

Key Questions

What performance does Mimic's FLUX-mimic achieve on soft-body kitting tasks?

FLUX-mimic reaches a 95% success rate on soft-body kitting, compared to 55% for π0.5. It uses video pre-training and shows that action performance scales directly with video model quality.

How does FLUX-mimic differ from traditional VLA approaches?

The method challenges the VLM-centric VLA paradigm by demonstrating direct scaling of robot actions with video model quality rather than vision-language models. This represents a shift toward video-action models for scalable robot learning.

What deployment advantages does FLUX-mimic offer?

Inference optimizations such as partial denoising, quantization, and RTC enable edge deployment on a single RTX 5090. It also supports 30-minute fine-tuning versus 30+ hours previously and is deployed at Audi.

Mimic's FLUX-mimic uses video pre-training to achieve 95% success on soft-body kitting vs 55% for π0.5. The method scales action performance directly with video model quality, challenging the VLM-centric VLA paradigm. Inference optimizations (partial denoising, quantization, RTC) make it edge-deployable on a single RTX 5090. New details: 30-minute fine-tuning vs 30+ hours, deployed at Audi. This is a concrete, novel breakthrough in scalable robot learning.

Sources (2)
Updated Jul 25, 2026