DeepSeek V4 Flash 'O731' Nearly Matches GPT-5.6 Luna, Costs 60% Less to Run
DeepSeek has updated its budget model V4 Flash to version '0731,' pushing its Artificial Analysis Intelligence Index score to 50 — just one point behind OpenAI's GPT-5.6 Luna — while costing an estimated 60 percent less per task. The MIT-licensed model keeps its 284B-parameter architecture but shows major gains in agentic benchmarks and token efficiency.
DeepSeek has released an updated version of its budget-tier model, V4 Flash "0731," that closes nearly all of the performance gap with OpenAI's GPT-5.6 Luna while undercutting it on price by roughly 60 percent per task.
According to the Artificial Analysis Intelligence Index, the new model scores 50 points, up ten points from the previous V4 Flash release in April 2026. That places it just one point behind GPT-5.6 Luna — even after OpenAI's own 80 percent price cut on that model. DeepSeek's cost advantage is driven partly by a 98 percent cache discount, well above the industry-standard 90 percent offered by most providers, according to Artificial Analysis.
What changed
The update improves scores across every category tested by Artificial Analysis, with the largest gains in agentic tasks. On GDPval, a benchmark measuring performance on complex, real-world office work, the model's Elo rating jumped from 1,189 to 1,559. DeepSeek also claims the model hallucinates less often than its predecessor and uses 12 percent fewer tokens per task, improving both speed and cost efficiency.
The underlying architecture is unchanged: 284 billion total parameters with 13 billion active per forward pass, and a one-million-token context window. This confirms the update is a fine-tuning or training refresh rather than a new base model.
Model weights are available under an MIT license on Hugging Face, continuing DeepSeek's practice of open-weighting its releases — a stark contrast to OpenAI and Anthropic's closed-model approach for comparable price-tier products.
Pricing and positioning
Exact per-token pricing for V4 Flash "0731" was not disclosed in the source data; the reported 60 percent cost advantage is a task-level comparison against GPT-5.6 Luna, factoring in DeepSeek's aggressive cache discount structure. Artificial Analysis lists the model as holding the top spot for price-to-performance ratio among currently tracked models in its index.
What this means
DeepSeek's update signals that the price war in budget-tier LLMs has not slowed down, even as OpenAI cuts GPT-5.6 Luna pricing by 80 percent. A one-point gap on the Intelligence Index at a fraction of the cost makes V4 Flash a strong default choice for high-volume, cost-sensitive workloads — particularly agentic and office-automation tasks, where the GDPval jump is largest.
The open-weight MIT license also matters: enterprises wary of vendor lock-in or data residency requirements can self-host a model that's now competitive with a top-tier closed alternative. For OpenAI and other closed-model providers, the pressure to justify pricing on budget tiers will likely intensify, since DeepSeek's cache-discount strategy suggests further cost reductions are structurally possible rather than one-off promotions.
Related Articles
OpenRouter Adds DeepSeek Flash Latest Alias With 1M-Token Context Window
OpenRouter has launched deepseek-flash-latest, a persistent endpoint that always points to the current DeepSeek Flash model. It offers a 1,049K token context window, text-and-image input, and pricing of $0.15 per 1M input tokens and $0.60 per 1M output tokens.
DeepSeek Releases V4.1-Flash: 552B MoE Model Cuts KV Cache to 890 Bytes Per Token
DeepSeek has released V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model supporting 1M-token context and activating only 8B parameters during prefill. The model uses a new Causal Encoder-Decoder architecture and Compressed Sparse Attention 2 to cut global KV cache to 890 bytes per token, roughly a quarter of its predecessor.
Meta Releases Muse Spark 1.3, Cheapest Model in Its Performance Class at $0.55 Per Task
Meta has released Muse Spark 1.3, its fourth model in five months, with an xhigh tier available now and a more powerful max tier in limited preview. The model improves sharply on agentic benchmarks and costs $0.55 per index task—cheaper than any rival at the same performance level—but still trails Claude Fable 5.1 on most tests.
Anthropic Releases Claude Fable 5.1 and Mythos 5.1, Cuts Cache Pricing 75% But Output Tokens Jump 70%
Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, claiming the top spot on Artificial Analysis's Intelligence Index at 66. Cache-read pricing dropped 75% to $0.25 per million tokens, but a 1.7x increase in output token usage pushes net per-task cost up 20%.
Comments
Loading...