DeepSeek Releases V4 Pro: 1.6T Parameter MoE Model with 1M Token Context at $1.74/M Input Tokens
DeepSeek has released V4 Pro, a Mixture-of-Experts model with 1.6 trillion total parameters and 49 billion activated parameters. The model supports a 1-million-token context window and costs $1.74 per million input tokens and $3.48 per million output tokens.
DeepSeek-V4-Pro — Quick Specs
DeepSeek Releases V4 Pro: 1.6T Parameter MoE Model with 1M Token Context
DeepSeek has released V4 Pro, a large-scale Mixture-of-Experts model with 1.6 trillion total parameters and 49 billion activated parameters, supporting a 1-million-token context window. The model is priced at $1.74 per million input tokens and $3.48 per million output tokens.
Architecture and Capabilities
According to DeepSeek, V4 Pro is built on the same architecture as DeepSeek V4 Flash and introduces a hybrid attention system designed for efficient long-context processing. The model supports multiple reasoning modes that allow users to balance speed and depth depending on task requirements.
The company claims the model delivers strong performance across knowledge, mathematics, and software engineering benchmarks, though specific benchmark scores have not been disclosed.
Target Use Cases
DeepSeek positions V4 Pro for complex workloads including:
- Full-codebase analysis
- Multi-step automation
- Large-scale information synthesis
- Advanced reasoning tasks
- Long-horizon agent workflows
The 1-million-token context window enables processing of entire codebases or lengthy documents in a single inference call.
Pricing and Availability
V4 Pro is available through OpenRouter as of April 24, 2026. At $1.74 per million input tokens, it sits in the mid-range pricing tier for frontier models. The 2:1 output-to-input pricing ratio ($3.48 vs $1.74) is standard for models with generation-heavy workloads.
Technical Details
The Mixture-of-Experts architecture activates 49 billion parameters per forward pass while maintaining 1.6 trillion total parameters. This approach aims to provide capabilities comparable to dense models of similar active parameter count while reducing computational costs.
OpenRouter integration includes support for DeepSeek's reasoning modes, with developers able to access step-by-step thinking processes through the reasoning_details array in API responses.
What This Means
V4 Pro represents DeepSeek's entry into the ultra-long-context market dominated by models like Anthropic's Claude and Google's Gemini. The 1M token context window and competitive pricing make it viable for enterprise use cases requiring analysis of large documents or codebases. However, without published benchmark scores, direct performance comparisons to established models remain unclear. The MoE architecture suggests DeepSeek is prioritizing inference efficiency alongside capability, a trend across Chinese AI labs competing with Western frontier model providers.
Related Articles
Poolside Releases Laguna S 2.1, an 8B-Active-Parameter Open Coding Model That Rivals Systems 20x Its Size
Poolside has released Laguna S 2.1, a mixture-of-experts coding model with 8 billion active parameters out of 118 billion total, its third coding model release in three months. The company claims it outperforms open-weight models 10 to 20 times its size on agentic coding benchmarks like Terminal-Bench 2.1 and DeepSWE.
Meituan launches LongCat 2.0: 1.6T parameter MoE model with 1M+ context window at $0.30 per 1M input tokens
Meituan has released LongCat 2.0, a sparse mixture-of-experts language model with 48 billion active parameters out of 1.6 trillion total. The model features a 1,049,000 token context window and costs $0.30 per 1M input tokens and $1.20 per 1M output tokens.
Moonshot AI Releases Kimi K3: 2.8T Parameter Open Model at $3/$15 Per Million Tokens
Moonshot AI has released Kimi K3, a 2.8 trillion parameter model with 1 million token context window and native multimodal input. The model ranks #1 in Frontend Code Arena and #9 in Text Arena, with pricing at $3 per million input tokens and $15 per million output tokens—comparable to Claude Sonnet 5 pricing while delivering performance the company claims is near Claude Opus 4.8 and GPT-5.5.
Thinking Machines releases Inkling: 975B-parameter MoE model with Apache 2.0 license, first major US open-weight multimo
Thinking Machines Lab released Inkling, a mixture-of-experts model with 975B total parameters and 41B active parameters, trained on 45 trillion tokens across text, images, audio, and video. The Apache 2.0-licensed model supports up to 1M context and debuts alongside Inkling-Small (276B-A12B), marking what observers call the strongest US-based open-weight release to date.
Comments
Loading...