model releaseMoonshot AI

Moonshot AI releases 2.8T parameter Kimi K3, pricing at $3/$15 per million tokens

TL;DR

Chinese AI lab Moonshot AI released Kimi K3, a 2.8 trillion parameter model priced at $3 per million input tokens and $15 per million output tokens. The model is currently available via API, with open weights promised by July 27, 2026. This represents the most expensive pricing from a Chinese AI lab to date, matching Anthropic's Claude Sonnet series.

2 min read
0

Kimi K3 — Quick Specs

Context window1049K tokens
Input$3/1M tokens
Output$15/1M tokens

Moonshot AI releases 2.8T parameter Kimi K3, pricing at $3/$15 per million tokens

Chinese AI lab Moonshot AI released Kimi K3, a 2.8 trillion parameter model priced at $3 per million input tokens and $15 per million output tokens. The model is currently available via API, with open weights promised by July 27, 2026.

Model specifications and pricing

Kimi K3 represents a significant scale increase from Moonshot's previous model, K2.6, which had 1 trillion parameters and was priced at $0.95 input/$4 output per million tokens. Moonshot claims K3 is their "most capable model to date" and describes it as the first "open 3T-class model," surpassing DeepSeek's 1.6T parameter v4 Pro.

The $3/$15 pricing matches Anthropic's Claude Sonnet series and makes K3 the most expensive model released by a Chinese AI lab. This represents more than a 3x price increase over K2.6.

Benchmark performance

According to Moonshot's self-reported benchmarks, K3 beats Claude Opus 4.8 max and GPT-5.5 high on most tasks, while trailing Claude Fable 5 and GPT-5.6 Sol. Third-party evaluator Artificial Analysis reports K3 achieved an Elo of 1547 on their private long-horizon knowledge work evaluation, placing it behind only Claude Fable 5. The model now leads Arena.ai's Frontend Code arena, surpassing Claude Fable 5.

Artificial Analysis reports K3's cost per task at $0.94, similar to GPT-5.6 Sol ($1.04), roughly half the price of Opus 4.8 ($1.80), though higher than open weight alternatives. The model uses 21% fewer output tokens than K2.6 according to the Artificial Analysis Intelligence Index.

Reasoning token usage

K3 currently offers only one reasoning effort level ("max"), which shows significant token consumption. A simple SVG generation prompt consumed 13,241 reasoning tokens to output 3,417 tokens of response, costing approximately $0.25 for that single request. The same prompt counted as 95 input tokens despite being only 10 words, suggesting an 85-token hidden system prompt.

The model includes vision capabilities, demonstrating strong performance on image description tasks.

What this means

Kimi K3's pricing strategy marks a shift for Chinese AI labs toward premium pricing tiers, matching Western competitors like Anthropic. The aggressive pricing suggests confidence in the model's capabilities, though the high reasoning token consumption at the "max" effort level could make practical applications expensive. The promised open weights release by July 27, 2026, will be significant given the model's 2.8T parameter size—making it the largest openly available model if released as planned. The model's strong performance on coding benchmarks and third-party evaluations suggests Chinese labs continue closing the gap with frontier Western models, though head-to-head comparisons remain challenging without standardized benchmarks.

Related Articles

model release

Tencent Open-Sources Hy4 Preview: 770B-Parameter MoE Model with 1M-Token Context

Tencent's Hy Team has open-sourced Hy4 preview, a 770-billion-parameter Mixture-of-Experts model with 49 billion activated parameters and a 1-million-token context window. The model is available under Apache 2.0 alongside an FP8-quantized variant, with Tencent claiming it beats GLM 5.3 and Kimi K3 on internal engineering evaluations.

model release

Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents

Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.

model release

Google Launches Gemini 3.5 Transcribe with 4.0% Word Error Rate Across 85 Languages

Google has released Gemini 3.5 Transcribe, a speech-to-text model that automatically detects 85 languages, removes filler words, and corrects misspoken phrases. The company claims a 4.0 percent word error rate for streaming audio and 70 percent lower latency than its predecessor, Chirp 3.

model release

Z.ai's GLM-5.3-Flash Matches Top Models at 7.5x Lower Cost, Runs Entirely on Chinese Chips

Z.ai released GLM-5.3-Flash, a 320-billion-parameter MoE model with an 18-billion active parameter count and a one-million-token context window. It nearly matches the larger GLM-5.3 on Artificial Analysis's Intelligence Index while costing roughly 7.5 times less per task, and it reportedly runs entirely on Chinese AI chips instead of Nvidia GPUs.

Comments

Loading...