DeepSeek V4 Flash 'O731' Nearly Matches GPT-5.6 Luna, Costs 60% Less to Run
DeepSeek has updated its budget model V4 Flash to version '0731,' pushing its Artificial Analysis Intelligence Index score to 50 — just one point behind OpenAI's GPT-5.6 Luna — while costing an estimated 60 percent less per task. The MIT-licensed model keeps its 284B-parameter architecture but shows major gains in agentic benchmarks and token efficiency.
DeepSeek has released an updated version of its budget-tier model, V4 Flash "0731," that closes nearly all of the performance gap with OpenAI's GPT-5.6 Luna while undercutting it on price by roughly 60 percent per task.
According to the Artificial Analysis Intelligence Index, the new model scores 50 points, up ten points from the previous V4 Flash release in April 2026. That places it just one point behind GPT-5.6 Luna — even after OpenAI's own 80 percent price cut on that model. DeepSeek's cost advantage is driven partly by a 98 percent cache discount, well above the industry-standard 90 percent offered by most providers, according to Artificial Analysis.
What changed
The update improves scores across every category tested by Artificial Analysis, with the largest gains in agentic tasks. On GDPval, a benchmark measuring performance on complex, real-world office work, the model's Elo rating jumped from 1,189 to 1,559. DeepSeek also claims the model hallucinates less often than its predecessor and uses 12 percent fewer tokens per task, improving both speed and cost efficiency.
The underlying architecture is unchanged: 284 billion total parameters with 13 billion active per forward pass, and a one-million-token context window. This confirms the update is a fine-tuning or training refresh rather than a new base model.
Model weights are available under an MIT license on Hugging Face, continuing DeepSeek's practice of open-weighting its releases — a stark contrast to OpenAI and Anthropic's closed-model approach for comparable price-tier products.
Pricing and positioning
Exact per-token pricing for V4 Flash "0731" was not disclosed in the source data; the reported 60 percent cost advantage is a task-level comparison against GPT-5.6 Luna, factoring in DeepSeek's aggressive cache discount structure. Artificial Analysis lists the model as holding the top spot for price-to-performance ratio among currently tracked models in its index.
What this means
DeepSeek's update signals that the price war in budget-tier LLMs has not slowed down, even as OpenAI cuts GPT-5.6 Luna pricing by 80 percent. A one-point gap on the Intelligence Index at a fraction of the cost makes V4 Flash a strong default choice for high-volume, cost-sensitive workloads — particularly agentic and office-automation tasks, where the GDPval jump is largest.
The open-weight MIT license also matters: enterprises wary of vendor lock-in or data residency requirements can self-host a model that's now competitive with a top-tier closed alternative. For OpenAI and other closed-model providers, the pressure to justify pricing on budget tiers will likely intensify, since DeepSeek's cache-discount strategy suggests further cost reductions are structurally possible rather than one-off promotions.
Related Articles
DeepSeek V4-Flash 0731 Update Jumps Terminal-Bench Score by 25.8 Points With No Architecture Change
DeepSeek released V4-Flash 0731, a post-training-only update to its API and open-weights model that lifted Terminal-Bench scores by 25.8 points without changing model architecture or parameter count. The update arrived alongside disclosed sandbox-escape incidents at OpenAI and Anthropic that renewed debate over eval infrastructure and open-weight safety.
OpenAI Refines GPT-5.6 Sol for ChatGPT, Unifies Instant/Thinking Modes, Makes Free Text Chat Unlimited
OpenAI is rolling out a ChatGPT-specific tuning of GPT-5.6 Sol that merges Instant and Thinking modes behind a new reasoning slider for Plus and Pro subscribers. Free users now get unlimited text chats with GPT-5.6 Luna and a new Think button.
DeepSeek Launches 'V4 Flash Latest' Alias with 1M+ Token Context on OpenRouter
DeepSeek has published a new routing endpoint, deepseek-v4-flash-latest, that always points to the newest model in its V4 Flash family. The endpoint offers a 1,049K token context window and pricing of $0.09/M input and $0.18/M output tokens via OpenRouter.
DeepSeek Releases V4-Flash-0731, a 284B-Parameter Model That Beats Its Own Larger Pro Variant on Agentic Benchmarks
DeepSeek has shipped the full release of DeepSeek-V4-Flash-0731, a 284B-parameter model that according to DeepSeek outperforms its own larger V4-Pro (Preview) on agentic and coding benchmarks. Unsloth has published quantized GGUF versions, with lossless 8-bit weights requiring 162GB of storage.
Comments
Loading...