model release

Meta Releases Muse Glimmer 30B, an Open-Weight Agentic Model for Consumer Hardware

TL;DR

Meta Superintelligence Labs has released Muse Glimmer 30B, a dense open-weight model distilled from its larger Muse Spark system and tuned for agentic workflows on consumer hardware. The model supports 131K context, image understanding, and over 100 languages at $0.30/$1.10 per 1M input/output tokens.

2 min read
0

Meta Muse Glimmer 30B — Quick Specs

Context window131K tokens
Input$0.3/1M tokens
Output$1.1/1M tokens

Meta Superintelligence Labs released Muse Glimmer 30B on August 9, 2026, a dense, open-weight multimodal model distilled from the company's larger Muse Spark series. The model is positioned for autonomous agent workloads that can run on consumer hardware rather than requiring data-center-scale infrastructure.

Specifications

Muse Glimmer 30B ships with a 131K-token context window and is priced at $0.30 per 1M input tokens and $1.10 per 1M output tokens through OpenRouter-listed providers, including Phala, DeepInfra, Fireworks, and Together. DeepInfra lists a slightly higher output rate of $1.20 per 1M tokens, and Fireworks and Together price output at $1.50 per 1M tokens, indicating some provider-side pricing variance around the base rate.

According to Meta, the model was distilled from Muse Spark and retains multi-step reasoning, tool-use reliability, and failure-recovery behavior intended for long-horizon agentic and coding tasks. It supports image understanding and multilingual input across more than 100 languages.

Benchmark results

Across provider-reported evaluations, Muse Glimmer 30B scored between 81.3% and 82.7% on GPQA Diamond and between 73.7% and 78.7% on TAU-Bench, depending on the serving provider. Parasail reported the top scores at 82.7% (GPQA Diamond) and 78.7% (TAU-Bench), while Together reported the lowest at 81.3% and 73.7% respectively. These are provider-side measurements published through OpenRouter rather than independent third-party verification, and the spread suggests output can vary based on serving infrastructure, quantization, or routing configuration.

Availability and reliability

Over a rolling three-day window, OpenRouter reported 100% uptime and 97.95% availability for the model, with per-provider latency ranging from 0.81 to 0.84 seconds at P50 and throughput up to 108 tokens per second on DeepInfra. Meta has released model weights, making Muse Glimmer 30B usable outside hosted API access, though local hardware requirements were not disclosed.

The release fits into a broader Muse family Meta has been building out, which includes Muse Spark 1.3 (a 1M-context reasoning model priced at $1.25/$4.25 per 1M tokens), a cheaper Muse Spark 1.3 Contributor tier at $0.10/$0.20 per 1M tokens, Muse Image for agentic image generation at $0.01 per image, and Muse Voice Transcribe 1.0 for speech-to-text at $0.00005 per second.

What this means

Muse Glimmer 30B fills a specific niche in Meta's lineup: a smaller, distilled, open-weight model meant to run agentic workflows without the compute footprint or cost of the flagship 1M-context Muse Spark models. At 30B dense parameters and a sub-$2 blended rate per million tokens, it competes directly with other open-weight agentic models aimed at developers who want local deployability alongside API access. The variance in benchmark scores across hosting providers is worth noting — it's a reminder that published benchmark numbers for an open-weight model can shift depending on who's serving it and how, something buyers should verify against their own provider before committing to a specific throughput or accuracy target.

Related Articles

model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

model release

NVIDIA Releases Nemotron 3 Diarization, an Open-Weight Speaker ID Model Supporting Up to 8 Speakers

NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model that determines "who spoke when" in audio, supporting both streaming and offline inference for up to eight speakers. The model achieves input buffer latency as low as 80 milliseconds and is available for commercial and non-commercial use.

model release

Fireworks Releases Ember-1, a Reasoning Model That Cuts Token Usage 40% Versus Its Kimi K3 Base

Fireworks Research has released Ember-1, a reasoning model built on Kimi K3 that produces shorter reasoning traces while claiming comparable output quality. The model offers a 1 million token context window at $3 per 1M input tokens and $15 per 1M output tokens.

model release

Anonymous Stealth Model "Space Bunny Alpha" Debuts on OpenRouter With 1M-Token Context, Free During Preview

A previously unknown AI provider has released Space Bunny Alpha, a stealth model on OpenRouter offering a 1M-token context window, adjustable reasoning effort, and multimodal input support. The model is free during its preview period, though its developer remains unnamed.

Comments

Loading...

Meta Muse Glimmer 30B: Open-Weight Agentic AI Model | TPS