Meta Opens Muse Glimmer, an Agent That Runs Locally

Muse Glimmer running on a graphics card set on a table inside a powered-down data center

Meta has released Muse Glimmer, a 30-billion-parameter agent model distributed under the Apache 2.0 licence and designed to run on a single consumer graphics card. It is the first release out of the Superintelligence Labs, and it puts the company back on ground it had walked away from months ago.

Key Takeaways

  • Muse Glimmer carries 30 billion parameters and downloads freely under Apache 2.0.
  • Quantised to 4 bits, it drops under 20 GB of memory and fits a 24 or 32 GB machine.
  • Meta places it ahead of Gemma4-31B and Qwen3.6-27B on several public test sets.

Have an AI Sum Up This Article

ChatGPT

A Model Built for Agents That Never Shut Down

Muse Glimmer does not land as one more generalist. Meta frames it as a model built for always-on local agents, with three promises spelled out in the technical sheet the company published alongside it: reliable tool calling, state that survives a restart, and memory the model manages on its own across sessions that run for hours.

That framing explains the sizing. Thirty billion parameters is small for a frontier model and large for a desktop machine. At full precision, a model that size demands more than 55 GB of memory, which puts it out of reach of nearly all consumer hardware.

The 4-bit quantised build brings that requirement under 20 GB. The full setup then fits inside 24 or 32 GB, which is precisely the shape of a recent high-end card or a well-specced Apple laptop. The threshold was engineered, not stumbled into.

The published scores follow the same agentic logic. MCP Atlas at 75.5 and DeepSearch QA at 74.6 on general agent work. SWE-Bench Verified at 76.0 and SWE-Bench Pro at 51.2 on coding. AIME 2026 at 94.7 and GPQA Diamond at 83.5 on reasoning. Meta positions the whole set ahead of Gemma4-31B and Qwen3.6-27B, its two direct rivals in the local category.

Speed got its own treatment. Meta pairs the model with a speculative decoding technique called DFlash, claiming a 3.1 times gain on an RTX 5090, 1.8 times on an M5 Max and 1.5 times on an M4 Max. On an agent grinding in the background all day, that multiplier matters more than a benchmark point.

The weights sit on the public model repository, methodology report included, with announced support for Ollama, LM Studio, llama.cpp, MLX, ExecuTorch, vLLM and SGLang. That is the entire tooling chain the local community already runs.


Muse Glimmer

What Muse Glimmer Changes for the People Installing It

The first effect shows up on the invoice. An agent running on its user’s own machine burns no metered tokens, and metered tokens are exactly the line item that blew up inside engineering teams this year. Usage cost moves from variable to fixed.

The contrast with the other end of the same company’s range is sharp. Meta separately sells an in-house agent billed at 200 dollars a month to clear entire work tasks. The same vendor now offers both business models, the premium subscription and the free download.

The second effect lands on reliability. Meta’s agents had a rough year on that front, to the point that the internal timeline was acknowledged as slower than planned. A model specialised in persistence and tool calling answers that criticism head on.

The trade-off is real and well documented. Four-bit quantisation is never free on output quality, and a 30-billion-parameter model stays far from the frontier on long, ambiguous tasks. Muse Glimmer targets repetition, not the one-off feat.

The pricing manoeuvre itself is familiar territory here. The company had already used price as its way into the market with Muse Spark 1.1. Going from undercutting to giving the thing away is the logical next step on that path.


More articles on Horizon


The Answer Now Owed by Google and Alibaba

By naming Gemma4-31B and Qwen3.6-27B as the benchmarks it beats, Meta points at its targets without hedging. Google and Alibaba had owned the slot for an open model small enough to sit on a personal machine, and that slot just got tighter.

The likeliest answer runs through agentic specialisation rather than raw size. A rival shipping a merely bigger model would miss the point: the Muse Glimmer argument is not brute power, it is holding up across long sessions with state that survives a restart.

The release fits a wider infrastructure play, where the company has started renting out its spare processing capacity. Giving away a model that runs on the user’s own hardware while selling raw capacity to enterprises are not contradictory moves.

Mark Zuckerberg paired the release with an essay in which he defends model distillation and argues for fewer regulatory constraints on American labs. The technical drop and the political position landed on the same day, which is never a scheduling accident at Meta.

That leaves the question of pace. An open-weights build of Muse Spark 1.2 is flagged as coming next, which would make two open releases within weeks after a long stretch of silence. If the pace holds, the pressure lands first on the Chinese labs that had the ground to themselves while Meta was away.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *