DeepSeek Launches V4.1 Flash: Low-Cost MoE Model Claims to Beat V4 Pro
DeepSeek has released V4.1 Flash, a sparse mixture-of-experts model priced at $0.30 per 1M input tokens and $1.20 per 1M output tokens with a 1 million token context window. DeepSeek claims the model exceeds the larger V4 Pro on performance, speed, and task completion time.
DeepSeek Releases V4.1 Flash
DeepSeek has released DeepSeek V4.1 Flash, the cost-efficient tier of its V4.1 model family. The model is a sparse mixture-of-experts (MoE) architecture and is now listed on OpenRouter with a 1 million token context window.
Pricing and Performance
V4.1 Flash is priced at $0.30 per 1M input tokens and $1.20 per 1M output tokens, with cache-read pricing at $0.006 per 1M tokens. According to OpenRouter's provider data, the model posts a P50 latency of 0.76 seconds and throughput of 134 tokens per second, among the fastest currently benchmarked on the platform for a model of this class.
DeepSeek claims that V4.1 Flash exceeds the larger V4 Pro model on performance, speed, and task completion time — an unusual positioning that puts a "Flash" tier ahead of the flagship it's nominally cheaper than. This claim has not been independently verified with third-party benchmark scores at time of writing; no MMLU, HumanEval, or other standard benchmark figures have been published alongside the release.
Intended Use Cases
DeepSeek positions V4.1 Flash for coding, reasoning, and agentic workflows, with particular emphasis on long-horizon tasks — multi-step processes that must run to completion across extended tool-use chains. The model supports a configurable reasoning-effort parameter, allowing developers to trade off latency and cost against depth of reasoning on a per-request basis.
Release Timing
OpenRouter's listing shows a release date of September 10, 2026. This date has not been corroborated elsewhere and should be treated as provisional pending confirmation from DeepSeek directly, as it falls outside the current publication window.
What This Means
V4.1 Flash continues DeepSeek's pattern of undercutting Western frontier-lab pricing while pushing large context windows — the 1M-token window plus sub-second latency at 134 tokens/sec is a genuinely competitive combination for agentic coding tools that need to hold large codebases or long conversation histories in context. The claim that a "Flash" tier beats the company's own "Pro" flagship is notable but unverified; without independent benchmark numbers (HumanEval, SWE-bench, MMLU, or similar), buyers should treat DeepSeek's performance claims as marketing until third-party evaluations appear. If the speed and cost figures hold up under real workloads, V4.1 Flash becomes a strong default choice for agentic and coding pipelines where token cost and end-to-end latency dominate the decision, rather than peak reasoning accuracy alone.
Related Articles
DeepSeek Releases V4.1-Flash: 552B MoE Model Cuts KV Cache to 890 Bytes Per Token
DeepSeek has released V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model supporting 1M-token context and activating only 8B parameters during prefill. The model uses a new Causal Encoder-Decoder architecture and Compressed Sparse Attention 2 to cut global KV cache to 890 bytes per token, roughly a quarter of its predecessor.
Alibaba Releases Qwen3.8 Max (0902), a 2.4-Trillion-Parameter MoE Model With 1M-Token Context
Alibaba's Qwen team released Qwen3.8 Max (0902), a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input. The snapshot is post-trained for coding, agentic workflows, and long-horizon task execution, priced at $2/$6 per 1M input/output tokens.
Alibaba Open-Sources Qwen3.8-2.4T-A95B, Its First Qwen-Max-Class Model With Public Weights
Alibaba's Qwen team released Qwen3.8-2.4T-A95B on August 12, 2026, the open-weight version of Qwen3.8-Max and the first Qwen-Max-class model made publicly available. The 2.4 trillion-parameter mixture-of-experts model activates only 95 billion parameters per token and supports context windows up to 1 million tokens.
Suno Releases v6 Music Model Family Trained on Licensed Data from Warner, BMG, Believe
Suno unveiled its v6 model family, trained on licensed data from Warner Music Group, BMG, and Believe, as the AI music startup continues fighting copyright lawsuits from Sony, Universal Music Group, and individual artists. The company plans to retire its older, non-licensed models entirely.
Comments
Loading...