model releaseDeepSeek

DeepSeek V4 Flash Released: 284B Parameter MoE Model with 1M Context Window at $0.14 per Million Tokens

TL;DR

DeepSeek has released V4 Flash, a Mixture-of-Experts model with 284B total parameters and 13B activated parameters per request. The model supports a 1,048,576-token context window and is priced at $0.14 per million input tokens and $0.28 per million output tokens.

2 min read
0

DeepSeek V4 Flash — Quick Specs

Context window1000K tokens
Input$0.098/1M tokens
Output$0.196/1M tokens

DeepSeek V4 Flash Released: 284B Parameter MoE Model with 1M Context Window at $0.14 per Million Tokens

DeepSeek has released V4 Flash, a Mixture-of-Experts model with 284B total parameters and 13B activated parameters per request. The model supports a 1,048,576-token context window and is priced at $0.14 per million input tokens and $0.28 per million output tokens.

Model Architecture and Capabilities

DeepSeek V4 Flash uses a sparse Mixture-of-Experts architecture that activates only 13B of its 284B total parameters for each inference request. According to DeepSeek, the model includes hybrid attention mechanisms designed for efficient long-context processing.

The model supports configurable reasoning modes, allowing it to show step-by-step thinking processes. DeepSeek claims the model maintains strong performance on reasoning and coding tasks despite its efficiency optimizations.

Pricing and Availability

The model is available through OpenRouter at:

  • Input: $0.14 per million tokens
  • Output: $0.28 per million tokens

These prices position V4 Flash as a cost-effective option for high-throughput workloads compared to larger models with similar context windows.

Target Use Cases

DeepSeek designed V4 Flash for applications requiring fast inference and high throughput, including:

  • Coding assistants
  • Chat systems
  • Agent workflows

The model's sparse activation pattern (activating only 4.6% of total parameters) enables faster inference speeds while attempting to preserve model quality.

Technical Details

Release date: April 24, 2026 (as listed on OpenRouter) Context window: 1,048,576 tokens Architecture: Sparse Mixture-of-Experts Reasoning support: Configurable reasoning modes with exposed thinking processes

What This Means

DeepSeek V4 Flash continues the trend of using sparse MoE architectures to deliver capable models at lower inference costs. The 13B activated parameter count per request allows for faster processing than dense models of similar capability, while the 1M token context window matches the extended context offerings from competitors like Anthropic and Google. The $0.14/$0.28 per million token pricing undercuts many competing models with similar context lengths, potentially making it attractive for high-volume production deployments where cost per token matters more than absolute peak performance.

Related Articles

model release

Alibaba Releases Qwen3.8 Max (0902), a 2.4-Trillion-Parameter MoE Model With 1M-Token Context

Alibaba's Qwen team released Qwen3.8 Max (0902), a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input. The snapshot is post-trained for coding, agentic workflows, and long-horizon task execution, priced at $2/$6 per 1M input/output tokens.

model release

InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters

InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.

model release

OpenAI Launches GPT-6 Astra With Half the Message Allowance of GPT-5.6 Sol

OpenAI has begun rolling out GPT-6 Astra to top-tier ChatGPT plans, the API, Azure, and AWS Bedrock. The model delivers roughly half the usage allowance of GPT-5.6 Sol across comparable plans, with Plus and Business users gaining access in the coming days.

model release

OpenAI Releases GPT-6 Astra, First Model to Cross 'Critical' Cybersecurity Threshold

OpenAI has begun rolling out GPT-6 Astra, the first model to reach the company's internal 'Critical' cybersecurity threshold. Access is being phased, with companies in OpenAI's Daybreak cybersecurity program getting priority following added safeguards after a prior model containment breach.

Comments

Loading...