DeepSeek-V4.1-Flash

DeepSeek🇨🇳 China
active
Context window1000K tokens
Input / 1M tokens$0.3
Output / 1M tokens$1.2

Version History

V4.1-Flashminor

DeepSeek-V4.1-Flash introduces a Causal Encoder-Decoder architecture and Compressed Sparse Attention 2 that cut global KV cache to 890 bytes per token, roughly one-quarter of DeepSeek-V4-Flash. The 552B-parameter multimodal MoE model supports 1M-token context and activates only 8B/16B parameters per token.

Benchmark Scores

Full leaderboard →
95.6%
DocVQA
79.4%
HumanEval

Coverage

model releaseDeepSeek

DeepSeek Releases V4.1-Flash: 552B MoE Model Cuts KV Cache to 890 Bytes Per Token

DeepSeek has released V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model supporting 1M-token context and activating only 8B parameters during prefill. The model uses a new Causal Encoder-Decoder architecture and Compressed Sparse Attention 2 to cut global KV cache to 890 bytes per token, roughly a quarter of its predecessor.

3 min read