Get the app
LLMs

Meta open-sourced a 30B agent model that fits your GPU

Muse Glimmer is Apache 2.0, runs offline on a 24GB card, and beats Qwen and Gemma on multi-step tool use.

Meta open-sourced a 30B agent model that fits your GPU

Meta Superintelligence Labs dropped Muse Glimmer on Hugging Face today under Apache 2.0 — a 30B dense multimodal model (2B vision encoder + 28B text decoder) distilled from the closed flagship Muse Spark. The pitch isn't raw intelligence, it's agentic competence: calling tools, writing and debugging code, reading screenshots, and recovering when a tool call fails mid-task.

The numbers back the framing. On MCP-Atlas, which measures multi-step tool orchestration, Glimmer scores 75.5 against 62.5 for Qwen3.6-27B and 54.2 for Gemma4-31B. It hits 51.2 on SWE-Bench Pro and 78.8 on CharXiv reasoning. More importantly, 4-bit quantization squeezes memory from ~55GB down to 18–20GB, so it lands inside a 24GB consumer card. Speculative decoding via DFlash gives a 3.1x speedup on an RTX 5090, 1.8x on an M5 Max. Day-one support ships for Ollama, LM Studio, llama.cpp, MLX, vLLM and SGLang.

The strategic read matters as much as the weights. Meta paired the release with a 6,500-word Zuckerberg essay pushing Washington to clear the path for open-source AI — plus a promise of open weights for Muse Spark 1.2 in coming weeks and a $1B fund for data-center host regions. After the Llama-license retreat, Meta is buying back open-source credibility with a permissive license and a real model.

Why it matters: an always-on personal agent that never phones home just became a 20GB download.

Sources

Primary: the company, paper or repository

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play