model releaseDeepSeek

DeepSeek V4 Pro launches with 1.6 trillion parameters, 1M token context at $0.145 per million input tokens

TL;DR

Chinese AI lab DeepSeek has released preview versions of DeepSeek V4 Flash and V4 Pro, mixture-of-experts models with 1 million token context windows. The V4 Pro has 1.6 trillion total parameters (49 billion active), making it the largest open-weight model available, while both models significantly undercut frontier model pricing.

2 min read
0

DeepSeek-V4-Pro — Quick Specs

Context window1000K tokens
Input$0.435/1M tokens
Output$0.87/1M tokens

DeepSeek V4 Pro launches with 1.6 trillion parameters, 1M token context at $0.145 per million input tokens

Chinese AI lab DeepSeek has released preview versions of DeepSeek V4 Flash and V4 Pro, both featuring 1 million token context windows and mixture-of-experts architectures that activate only a subset of parameters per task to reduce inference costs.

Model specifications

DeepSeek V4 Pro contains 1.6 trillion total parameters with 49 billion active parameters, making it the largest open-weight model available. This exceeds Moonshot AI's Kimi K 2.6 (1.1 trillion parameters), MiniMax's M1 (456 billion), and more than doubles DeepSeek's previous V3.2 model (671 billion parameters).

The smaller V4 Flash has 284 billion total parameters with 13 billion active parameters.

Both models support text only, lacking the multimodal capabilities (audio, video, image) found in many closed-source frontier models.

Performance claims

According to DeepSeek, the V4-Pro-Max model outperforms open-source peers across reasoning benchmarks and surpasses OpenAI's GPT-5.2 and Gemini 3.0 Pro on some tasks. The company claims both V4 models perform comparably to GPT-5.4 on coding competition benchmarks.

However, DeepSeek acknowledges the models lag behind frontier models in knowledge tests, specifically OpenAI's GPT-5.4 and Google's Gemini 3.1 Pro. The lab states this represents "a developmental trajectory that trails state-of-the-art frontier models by approximately 3 to 6 months."

Pricing

DeepSeek V4 Flash: $0.14 per million input tokens, $0.28 per million output tokens — undercutting GPT-5.4 Nano, Gemini 3.1 Flash, GPT-5.4 Mini, and Claude Haiku 4.5.

DeepSeek V4 Pro: $0.145 per million input tokens, $3.48 per million output tokens — undercutting Gemini 3.1 Pro, GPT-5.5, Claude Opus 4.7, and GPT-5.4.

Timing and controversy

The launch follows U.S. accusations that China has stolen American AI labs' intellectual property using thousands of proxy accounts. DeepSeek has been specifically accused by Anthropic and OpenAI of "distilling" (copying) their AI models.

What this means

DeepSeek V4 Pro represents the largest open-weight model by parameter count, though its mixture-of-experts architecture means only 3% of parameters are active during inference. The 1 million token context window and aggressive pricing position these models as cost-effective alternatives for developers working with large codebases or documents, despite trailing frontier models by several months in knowledge capabilities and lacking multimodal support. The text-only limitation and acknowledged performance gap suggest DeepSeek is prioritizing efficiency and cost over feature parity with leading closed-source models.

Related Articles

model release

Poolside Releases Laguna S 2.1, an 8B-Active-Parameter Open Coding Model That Rivals Systems 20x Its Size

Poolside has released Laguna S 2.1, a mixture-of-experts coding model with 8 billion active parameters out of 118 billion total, its third coding model release in three months. The company claims it outperforms open-weight models 10 to 20 times its size on agentic coding benchmarks like Terminal-Bench 2.1 and DeepSWE.

model release

Alibaba releases Qwen 3.8, a 2.4 trillion parameter open-weight model claiming second place behind Fable 5

Alibaba has released Qwen 3.8, a 2.4 trillion parameter open-weight model that the company claims trails only Fable 5. The multimodal model processes images, videos, and documents, with a preview available through Alibaba's platforms at 10 percent of standard pricing.

model release

Moonshot AI Releases Kimi K3: 2.8T Parameter Open Model at $3/$15 Per Million Tokens

Moonshot AI has released Kimi K3, a 2.8 trillion parameter model with 1 million token context window and native multimodal input. The model ranks #1 in Frontend Code Arena and #9 in Text Arena, with pricing at $3 per million input tokens and $15 per million output tokens—comparable to Claude Sonnet 5 pricing while delivering performance the company claims is near Claude Opus 4.8 and GPT-5.5.

model release

Thinking Machines releases Inkling: 975B-parameter MoE model with Apache 2.0 license, first major US open-weight multimo

Thinking Machines Lab released Inkling, a mixture-of-experts model with 975B total parameters and 41B active parameters, trained on 45 trillion tokens across text, images, audio, and video. The Apache 2.0-licensed model supports up to 1M context and debuts alongside Inkling-Small (276B-A12B), marking what observers call the strongest US-based open-weight release to date.

Comments

Loading...