Meta Releases Muse Glimmer, a 30B Open-Weight Agent Model That Runs on a Single RTX 3090
Meta released Muse Glimmer, a 30B-parameter open-weight model under Apache 2.0 built for always-on local agents, alongside a promise to release Muse Spark 1.2 weights soon. The model runs on a single RTX 3090 and scores 35 on Artificial Analysis's Intelligence Index.
Meta released Muse Glimmer, a 30B-parameter open-weight model licensed under Apache 2.0, marking the company's first genuine frontier-adjacent open release since ramping activity around its Meta Superintelligence Labs (MSL) division. The announcement, posted by AI at Meta and amplified by Mark Zuckerberg and Alexandr Wang, arrived alongside a promise to release Muse Spark 1.2 weights "soon."
What Muse Glimmer Is
Glimmer is a dense, multimodal model explicitly optimized for local, always-on agent workflows rather than general chat. According to Meta, the model is designed for long-horizon agent loops and tool use, and can run entirely on consumer hardware — Meta states it fits on a single RTX 3090.
Meta's serving stack quantizes the model to under 20GB and pairs it with a lightweight speculative-decoding component called DFlash, which Meta claims delivers "fluid" local interaction. Community technical analysis, including from researchers Elie Bakouch and nrehiew, points to architectural choices resembling Gemma 4's hybrid attention with scale-free QK normalization, a deeper vision stack, and longer sliding-window attention (SWA). Notably, Glimmer was reportedly logit-distilled from the larger Muse Spark model and trained from the start on agentic traces — a departure from the conventional base-model-then-post-train pipeline.
Benchmark Standing
Third-party evaluation from Artificial Analysis places Muse Glimmer at 35 on its Intelligence Index, just behind Qwen3.6-27B (38) and roughly on par with Kimi K2.5 (36). On openness, Glimmer scores 44 on Artificial Analysis's Openness Index, reflecting its Apache 2.0 license and full weight release. Deployment footprint: approximately 60GB in BF16 and roughly 18GB at 4-bit quantization, with a 128K token context window and a memory-efficient hybrid attention design suited to single-node deployment. Pricing is not applicable in the traditional sense since the weights are open, though pricing for any Meta-hosted API access has not yet been disclosed.
The Broader Context
The release lands one year after Zuckerberg's original "Personal Superintelligence" essay and coincides with a companion essay reiterating Meta's stated mission to keep advanced AI capability in the hands of individuals rather than concentrated among institutions. Zuckerberg's essay lays out predictions around personal agents, entrepreneurial tooling, and free or affordable access to advanced AI, alongside stated risks including labor market disruption, AI misuse in cybersecurity and bioterrorism, and the dangers of centralized superintelligence achieved through recursive self-improvement.
What This Means
Muse Glimmer is a meaningful, if modest, re-entry for Meta into competitive open-weight releases after a period of cautious, incremental launches (Dreamer, Muse Spark, Muse Code). At 35 on the Intelligence Index, Glimmer doesn't lead its size class outright — Qwen3.6-27B edges it out — but its combination of small footprint, agent-specific training, and Apache 2.0 licensing targets a different use case: self-hosted, local-first agents rather than raw benchmark supremacy. The real test comes with Muse Spark 1.2's promised release, which will reveal whether Meta intends to compete at the larger end of the open-weight frontier or continue positioning smaller, deployment-optimized models as its primary open contribution.
Related Articles
Black Forest Labs Releases FLUX 3 Action, a 7B Open-Weights World Action Model, Claims Top RoboLab Benchmark Score
Black Forest Labs has released FLUX 3 Action, a 7B parameter open-weights World Action Model. The company claims it achieves first place on the RoboLab benchmark, though independent verification is pending.
Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context
Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.
Qwen3.8 Omni Flash: Alibaba's First Agentic Omni-Modal Model Adds Native Audio-Video Understanding, 1M Context
Alibaba's Qwen team has released Qwen3.8 Omni Flash, described as the first Qwen model built around agentic capabilities with native audio-video understanding. It ships with a 1M-token context window and support for two- and four-channel spatial audio.
Black Forest Labs Releases FLUX 3 Action, a 7B-Parameter Open Robotics Model
Black Forest Labs has released FLUX 3 Action, an open-weight robotics model built on its FLUX 3 multimodal foundation. The 7-billion-parameter model reads multi-camera video feeds and predicts what a robot should do next, claiming a record success rate on the RoboLab-120 leaderboard while running nearly 4x faster than the previous best open model.
Comments
Loading...