New open vision models target long-context and embodied agents
Apple's LensVLM-9B appeared on Hugging Face on September 21, 2026, using image-to-text compression and selective high-resolution decoding for long-context vision, but its benchmarks, license, code, and demos are still unclear. Related Qwen3.8-27B robotics work reports 30/45 generalization trials across nine manipulation tasks and a 50% repetition-speed gain, while Microsoft's Rho collection adds physical-AI checkpoints with limited documentation.
Sources (4)
Updated Sep 24, 2026