Multimodal and Video Generation Advances: EchoWM, Alibaba Wan3.0, DeepSeek Flash Vision, and New Infrastructure
EchoWM: open omnimodal world model generating 720p video, sound, music, speech with trajectory control—strong multimodal advance. Alibaba launches Wan3.0 AI video model with 30s video, multi-source input, per-second pricing. DeepSeek Flash Vision emerges as new open-source multimodal king—10x cheaper than Kimi K3, beats GLM 5.3, excels at agentic coding with images. Today's reading adds: WeMM-Embedding (WeChat multimodal embedding family, SOTA 80.6 on MMEB-v2, open-sourced) and LAION-BVD (10M-hour open video dataset for multimodal pre-training). These signal rapid progress in multimodal generation, vision-language models, and infrastructure for builders.