DeepSeek Releases V4 Pro: 1.6T Parameter MoE Model with 1M Token Context at $1.74/M Input Tokens
DeepSeek has released V4 Pro, a Mixture-of-Experts model with 1.6 trillion total parameters and 49 billion activated parameters. The model supports a 1-million-token context window and costs $1.74 per million input tokens and $3.48 per million output tokens.
DeepSeek V4 Pro — Quick Specs
DeepSeek Releases V4 Pro: 1.6T Parameter MoE Model with 1M Token Context
DeepSeek has released V4 Pro, a large-scale Mixture-of-Experts model with 1.6 trillion total parameters and 49 billion activated parameters, supporting a 1-million-token context window. The model is priced at $1.74 per million input tokens and $3.48 per million output tokens.
Architecture and Capabilities
According to DeepSeek, V4 Pro is built on the same architecture as DeepSeek V4 Flash and introduces a hybrid attention system designed for efficient long-context processing. The model supports multiple reasoning modes that allow users to balance speed and depth depending on task requirements.
The company claims the model delivers strong performance across knowledge, mathematics, and software engineering benchmarks, though specific benchmark scores have not been disclosed.
Target Use Cases
DeepSeek positions V4 Pro for complex workloads including:
- Full-codebase analysis
- Multi-step automation
- Large-scale information synthesis
- Advanced reasoning tasks
- Long-horizon agent workflows
The 1-million-token context window enables processing of entire codebases or lengthy documents in a single inference call.
Pricing and Availability
V4 Pro is available through OpenRouter as of April 24, 2026. At $1.74 per million input tokens, it sits in the mid-range pricing tier for frontier models. The 2:1 output-to-input pricing ratio ($3.48 vs $1.74) is standard for models with generation-heavy workloads.
Technical Details
The Mixture-of-Experts architecture activates 49 billion parameters per forward pass while maintaining 1.6 trillion total parameters. This approach aims to provide capabilities comparable to dense models of similar active parameter count while reducing computational costs.
OpenRouter integration includes support for DeepSeek's reasoning modes, with developers able to access step-by-step thinking processes through the reasoning_details array in API responses.
What This Means
V4 Pro represents DeepSeek's entry into the ultra-long-context market dominated by models like Anthropic's Claude and Google's Gemini. The 1M token context window and competitive pricing make it viable for enterprise use cases requiring analysis of large documents or codebases. However, without published benchmark scores, direct performance comparisons to established models remain unclear. The MoE architecture suggests DeepSeek is prioritizing inference efficiency alongside capability, a trend across Chinese AI labs competing with Western frontier model providers.
Related Articles
Alibaba Releases Qwen3.8 Max (0902), a 2.4-Trillion-Parameter MoE Model With 1M-Token Context
Alibaba's Qwen team released Qwen3.8 Max (0902), a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input. The snapshot is post-trained for coding, agentic workflows, and long-horizon task execution, priced at $2/$6 per 1M input/output tokens.
OpenAI Releases Astra, Claims New Flagship Model Beats Rivals on Coding and Cybersecurity Benchmarks
OpenAI released Astra on Thursday, calling it its most capable and most aligned model yet. The model uses a reasoning technique called 'opaque recurrence' that critics say reduces visibility into its chain of thought.
InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters
InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.
Meta Releases Muse Spark 1.3, a Free Multimodal Reasoning Model with 1M-Token Context
Meta has released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window, listed as free on OpenRouter. The model targets long-running agentic, multi-agent, and coding workflows, though audio input support remains incomplete.
Comments
Loading...