Continuous unified multimodal generation
Multimodal Flow applies flow matching directly to ordered text-image hyperchunks, while PixelDense uses dense geometric and semantic alignment to improve pixel diffusion for generation, reconstruction, and editing. Reka's preliminary Rho-1 preview adds a unified text-image-video-action model with next-token and flow-matching objectives, but lacks standard benchmark and robotics evidence. Evaluation breadth and compute-efficiency comparisons remain unresolved.
Sources (2)
Updated Oct 6, 2026