Meta Muse Glimmer 30B: Open-weight agentic model for local deployment
Meta released Muse Glimmer (30B, Apache 2.0), a distilled agentic model optimized for local deployment. It beats Qwen3.6-27B and Gemma4-31B on benchmarks, quantizes to under 20GB, and supports dFlash speculative decoding. Day-0 llama.cpp support with official GGUF quant and AMD performance numbers (24 t/s on Ryzen, 53 t/s on Radeon) make it immediately runnable in 32-64GB setups. LM Studio and Lemonade integration further simplify deployment. Official confirmation from Yann LeCun.
Sources (2)
Updated Aug 12, 2026