开源多模态模型与视频任务生态持续扩张SenseNova U1.5 Lite、WeMM-Embedding、MOSS-VL 和 LAION-BVD 分别覆盖图像生成/理解、多模态检索表征、长上下文视频理解和大规模视频预训练数据。MMLVE 等多镜头视频编辑任务与 Aphanta 的视觉中间表示诊断,进一步扩展了视频和视觉推理的评测术语体系。Sources (3)@_akhaliq: VGI-Bench Probing Visual Intelligence in Video Generation Models paper: https://t.co/nHII38Xnr2 ht...x.comThinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoningarxiv.orgInternVL - OpenGVLab 推出的多模态大模型ai-bot.cnUpdated Aug 28, 2026