changelog

Meta Releases Muse Spark 1.3, Cheapest Model in Its Performance Class at $0.55 Per Task

TL;DR

Meta has released Muse Spark 1.3, its fourth model in five months, with an xhigh tier available now and a more powerful max tier in limited preview. The model improves sharply on agentic benchmarks and costs $0.55 per index task—cheaper than any rival at the same performance level—but still trails Claude Fable 5.1 on most tests.

3 min read
0

Meta has released Muse Spark 1.3 through Muse Code and the Meta Model API, its fourth model release in five months following the series' April launch, version 1.1 in July, and 1.2 in August.

The update ships in two tiers. The xhigh tier is available immediately. The more compute-heavy max tier is running only as a limited partner preview while Meta completes further safety testing, according to Artificial Analysis.

Pricing holds steady, cost per task rises

Meta kept per-token pricing unchanged at $1.25 per million input tokens and $4.25 per million output tokens. But because Muse Spark 1.3 uses more reasoning tokens than its predecessor, the effective cost of running one Artificial Analysis Intelligence Index task rose to $0.55, up from $0.40 for version 1.2.

Even so, no model scoring 59 points or higher on the index is cheaper. Rivals at the same performance tier charge between $0.94 and $1.23 per task. Meta has not disclosed pricing for the max variant.

Intelligence Index gains concentrate in agentic tests

On the Intelligence Index, max scores 62 and xhigh scores 61, up from 57 in August's version 1.2 and 53 in July's version 1.1. The improvement traces almost entirely to three heavily weighted benchmarks: GDPval-AA v2 (20% of the index), Terminal-Bench 2.1 (16%), and τ³-Bench Banking (14%).

On τ³-Bench Banking, which tests agents operating tools in a simulated banking scenario, max reaches 52 percent — currently the highest score of any model, according to Artificial Analysis. The available xhigh tier hits 47 percent, tying Claude Fable 5.1 (max) and GLM-5.3-Flash rather than leading outright. Version 1.2 scored 35 percent on the same test.

On Terminal-Bench 2.1, which measures terminal-based coding, xhigh climbs from 80 to 85 percent and max reaches 86 percent. Claude Fable 5.1 still leads at 91.4 percent (max), 91.0 percent (xhigh), and 89.9 percent (high).

On GDPval-AA v2 — a benchmark spanning 220 real-world professional tasks, scaled so human expert performance equals 1,000 — Meta improved from 1,615 to 1,709 (xhigh) and 1,754 (max). Claude Fable 5.1 (max) sits at 1,853. Meta's max variant achieves its edge over xhigh by burning 62 percent more reasoning tokens.

Mixed results outside agentic benchmarks

On GPQA Diamond, a set of expert-level science questions, Muse Spark 1.3 rose from 90 to 94 percent — placing in the top group but still behind Gemini 3.8 Flash (high) at 95.3 percent and Grok 4.6 (high) at 94.9 percent.

On CritPt, a research-physics benchmark, the model jumped from 18 to 26 percent, trailing GPT-5.6 Sol (max) at 32.3 percent and Claude Fable 5.1 (xhigh) at 31.1 percent.

Two scores declined relative to version 1.2. AA-LCR fell from 83 to 79 percent, and factual accuracy on AA-Omniscience dropped by up to three points — Meta attributes this to the model declining to answer more often when uncertain.

Meta says larger models and an open-weights release of the Muse Spark series are planned but has not given a timeline.

What this means

Muse Spark 1.3 does not lead the field on raw capability — Claude Fable 5.1 still tops most individual benchmarks, sometimes by wide margins on tests like GDPval-AA v2 and Terminal-Bench 2.1. Meta's play is cost efficiency: at $0.55 per index task, it undercuts every model in its performance bracket by 40 percent or more. For enterprises running high-volume agentic workloads — the exact use case the index's heaviest-weighted tests measure — that price gap may matter more than a few points of benchmark headroom. The declining factual-accuracy score is worth watching, since a model that answers less to avoid errors trades one failure mode for another.

Related Articles

changelog

Anthropic Releases Claude Fable 5.1 and Mythos 5.1, Cuts Cache Pricing 75% But Output Tokens Jump 70%

Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, claiming the top spot on Artificial Analysis's Intelligence Index at 66. Cache-read pricing dropped 75% to $0.25 per million tokens, but a 1.7x increase in output token usage pushes net per-task cost up 20%.

changelog

Google Launches Gemini 3.8 Flash, Warns It May Use More Tokens Despite Unchanged Pricing

Google released Gemini 3.8 Flash just weeks after Gemini 3.7 Flash, keeping the same per-token pricing of $0.75/$3.75 per million input/output tokens but warning it may consume more tokens overall. The model also ships with a cyber-focused variant restricted to a new government partner program called Fairwind.

changelog

Google Rolls Out Gemini 3.8 Flash, Third Flash Update in Three Months

Google has released Gemini 3.8 Flash, the third Flash-tier update in three months, arriving just three weeks after Gemini 3.7 Flash. The model is live now in the Gemini app, Google Antigravity, and AI Studio with introductory pricing of $0.75/1M input and $3.75/1M output tokens.

changelog

Google Adds Agent-Based Video Analysis to Gemini Flash, Cutting Token Usage by Up to 88 Percent

Google is rolling out agent-based video analysis for Gemini Flash models that dynamically searches footage instead of scanning frame by frame. Google claims the approach cuts token usage by up to 88 percent and costs by 66 percent while improving accuracy, with no added API fee.

Comments

Loading...