model release

Fireworks Releases Ember-1, a Reasoning Model That Cuts Token Usage 40% Versus Its Kimi K3 Base

TL;DR

Fireworks Research has released Ember-1, a reasoning model built on Kimi K3 that produces shorter reasoning traces while claiming comparable output quality. The model offers a 1 million token context window at $3 per 1M input tokens and $15 per 1M output tokens.

2 min read
0

Ember-1 — Quick Specs

Context window1000K tokens
Input$3/1M tokens
Output$15/1M tokens

Fireworks Research has released Ember-1, a reasoning model built on top of Moonshot AI's Kimi K3, designed to reduce the token overhead typically associated with chain-of-thought reasoning models.

According to Fireworks, Ember-1 uses roughly 40% fewer tokens than its Kimi K3 base model while maintaining comparable quality across the company's internal evaluations. The model targets coding, knowledge work, and agentic workflows — use cases where reasoning cost and latency directly affect production economics.

Specifications

Ember-1 ships with a 1 million token context window, positioning it among the largest-context models currently available through OpenRouter. Pricing is set at $3 per 1M input tokens and $15 per 1M output tokens, with cached input reads priced at $0.30 per 1M tokens.

On Fireworks' own infrastructure, the model posts a P50 latency of 0.57 seconds and throughput of 45 tokens per second — figures reported as best-in-class among providers currently serving the model. OpenRouter lists the model's availability at 98.57% over the trailing 24-hour window, with a three-day uptime history currently on record.

The model was listed with a release date of September 24, 2026, according to OpenRouter's model page.

Context: built on Kimi K3

Ember-1 is not a from-scratch model. Fireworks describes it as derived from Kimi K3, the reasoning model from Moonshot AI. Fireworks Research's contribution is a post-training or distillation process aimed at compressing reasoning traces — the intermediate "thinking" tokens a reasoning model generates before producing a final answer — without materially degrading output quality, according to Fireworks' evaluations. No independent third-party benchmark scores were provided in the source material, so the 40% reduction figure and "comparable quality" claim should be treated as Fireworks' own assessment pending outside verification.

No parameter count, training cutoff date, or public benchmark suite results (e.g., MMLU, HumanEval, GPQA) were disclosed for Ember-1.

What this means

Token efficiency has become a competitive axis for reasoning models as enterprises scale agentic workloads that make repeated calls with long reasoning chains. A 40% reduction in reasoning tokens, if it holds up under independent testing, would meaningfully lower per-task cost and latency for high-volume coding and agent use cases — often mattering more in production than raw benchmark scores.

That said, Fireworks is positioning Ember-1 as an efficiency layer on top of an existing open model (Kimi K3) rather than a new foundation model. This reflects a broader trend of infrastructure providers competing on serving efficiency and post-training optimization rather than pretraining from scratch. Buyers evaluating Ember-1 should request task-specific benchmarks rather than relying solely on the vendor's aggregate efficiency claim, since reasoning-trace compression can affect accuracy differently across domains like math, coding, and open-ended agentic planning.

Related Articles

model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

model release

Anonymous Stealth Model "Space Bunny Alpha" Debuts on OpenRouter With 1M-Token Context, Free During Preview

A previously unknown AI provider has released Space Bunny Alpha, a stealth model on OpenRouter offering a 1M-token context window, adjustable reasoning effort, and multimodal input support. The model is free during its preview period, though its developer remains unnamed.

model release

Aion Labs Launches Aion 3.5 Mini, a $0.70/M-Token Roleplaying Model with 262K Context

Aion Labs has released Aion 3.5 Mini, a lower-cost version of its multi-model roleplaying system Aion 3.5. Built on the GLM model family, it offers a 262K token context window at $0.70 per 1M input tokens and $1.40 per 1M output tokens.

model release

AionLabs Launches Aion 3.5, a Multi-Model Storytelling System Built on GLM

AionLabs has released Aion 3.5, a collaborative multi-model system for roleplaying and storytelling built on the GLM model family. It offers a 262K token context window at $3 per 1M input tokens and $6 per 1M output tokens.

Comments

Loading...