Meta open-sourced a 30B agent model that fits your GPU
Muse Glimmer is Apache 2.0, runs offline on a 24GB card, and beats Qwen and Gemma on multi-step tool use.

Meta Superintelligence Labs dropped Muse Glimmer on Hugging Face today under Apache 2.0 — a 30B dense multimodal model (2B vision encoder + 28B text decoder) distilled from the closed flagship Muse Spark. The pitch isn't raw intelligence, it's agentic competence: calling tools, writing and debugging code, reading screenshots, and recovering when a tool call fails mid-task.
The numbers back the framing. On MCP-Atlas, which measures multi-step tool orchestration, Glimmer scores 75.5 against 62.5 for Qwen3.6-27B and 54.2 for Gemma4-31B. It hits 51.2 on SWE-Bench Pro and 78.8 on CharXiv reasoning. More importantly, 4-bit quantization squeezes memory from ~55GB down to 18–20GB, so it lands inside a 24GB consumer card. Speculative decoding via DFlash gives a 3.1x speedup on an RTX 5090, 1.8x on an M5 Max. Day-one support ships for Ollama, LM Studio, llama.cpp, MLX, vLLM and SGLang.
The strategic read matters as much as the weights. Meta paired the release with a 6,500-word Zuckerberg essay pushing Washington to clear the path for open-source AI — plus a promise of open weights for Muse Spark 1.2 in coming weeks and a $1B fund for data-center host regions. After the Llama-license retreat, Meta is buying back open-source credibility with a permissive license and a real model.
Why it matters: an always-on personal agent that never phones home just became a 20GB download.
Sources
- Meta returns to open source with Muse Glimmer, an Apache 2.0 licensed 30B parameter AI model optimized for agents venturebeat.com
- Meta is back with Muse Glimmer: local, agentic, multimodal, and open source huggingface.co
Written by an AI pipeline from the sources above. How it works.
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.