DeepSeek V4 Flash 'O731' Nearly Matches GPT-5.6 Luna, Costs 60% Less to Run
DeepSeek has updated its budget model V4 Flash to version '0731,' pushing its Artificial Analysis Intelligence Index score to 50 — just one point behind OpenAI's GPT-5.6 Luna — while costing an estimated 60 percent less per task. The MIT-licensed model keeps its 284B-parameter architecture but shows major gains in agentic benchmarks and token efficiency.
DeepSeek has released an updated version of its budget-tier model, V4 Flash "0731," that closes nearly all of the performance gap with OpenAI's GPT-5.6 Luna while undercutting it on price by roughly 60 percent per task.
According to the Artificial Analysis Intelligence Index, the new model scores 50 points, up ten points from the previous V4 Flash release in April 2026. That places it just one point behind GPT-5.6 Luna — even after OpenAI's own 80 percent price cut on that model. DeepSeek's cost advantage is driven partly by a 98 percent cache discount, well above the industry-standard 90 percent offered by most providers, according to Artificial Analysis.
What changed
The update improves scores across every category tested by Artificial Analysis, with the largest gains in agentic tasks. On GDPval, a benchmark measuring performance on complex, real-world office work, the model's Elo rating jumped from 1,189 to 1,559. DeepSeek also claims the model hallucinates less often than its predecessor and uses 12 percent fewer tokens per task, improving both speed and cost efficiency.
The underlying architecture is unchanged: 284 billion total parameters with 13 billion active per forward pass, and a one-million-token context window. This confirms the update is a fine-tuning or training refresh rather than a new base model.
Model weights are available under an MIT license on Hugging Face, continuing DeepSeek's practice of open-weighting its releases — a stark contrast to OpenAI and Anthropic's closed-model approach for comparable price-tier products.
Pricing and positioning
Exact per-token pricing for V4 Flash "0731" was not disclosed in the source data; the reported 60 percent cost advantage is a task-level comparison against GPT-5.6 Luna, factoring in DeepSeek's aggressive cache discount structure. Artificial Analysis lists the model as holding the top spot for price-to-performance ratio among currently tracked models in its index.
What this means
DeepSeek's update signals that the price war in budget-tier LLMs has not slowed down, even as OpenAI cuts GPT-5.6 Luna pricing by 80 percent. A one-point gap on the Intelligence Index at a fraction of the cost makes V4 Flash a strong default choice for high-volume, cost-sensitive workloads — particularly agentic and office-automation tasks, where the GDPval jump is largest.
The open-weight MIT license also matters: enterprises wary of vendor lock-in or data residency requirements can self-host a model that's now competitive with a top-tier closed alternative. For OpenAI and other closed-model providers, the pressure to justify pricing on budget tiers will likely intensify, since DeepSeek's cache-discount strategy suggests further cost reductions are structurally possible rather than one-off promotions.
Related Articles
DeepSeek Releases V4-Flash-0731, a 284B-Parameter Model That Beats Its Own Larger Pro Variant on Agentic Benchmarks
DeepSeek has shipped the full release of DeepSeek-V4-Flash-0731, a 284B-parameter model that according to DeepSeek outperforms its own larger V4-Pro (Preview) on agentic and coding benchmarks. Unsloth has published quantized GGUF versions, with lossless 8-bit weights requiring 162GB of storage.
OpenAI Cuts GPT-5.6 Prices Up to 80%, Says Model's Own Self-Optimization Work Drove the Savings
OpenAI cut GPT-5.6 Luna pricing by 80% to $0.20/$1.20 per million tokens and GPT-5.6 Terra by 20% to $2/$12, while adding a 2.5x-faster mode for Sol at double the price. The company says GPT-5.6 itself rewrote production inference kernels and tuned its own speculative decoding pipeline to enable the cuts.
OpenAI Slashes GPT-5.6 Luna Pricing by 80%, Cuts Terra by 20%
OpenAI cut GPT-5.6 Luna pricing by 80 percent to $0.20 per million input tokens and $1.20 per million output tokens, while Terra dropped 20 percent to $2/$12. The company attributes the cuts to infrastructure efficiency gains and mounting price competition, particularly from Chinese providers.
OpenAI Cuts GPT-5.6 Terra Price 20%, Luna Price 80% Across API and ChatGPT
OpenAI is cutting API prices for its GPT-5.6 Terra and Luna models by 20% and 80%, respectively, compared to prices set earlier this month. The company says the lower costs are also reflected in usage limits for ChatGPT Work and Codex subscribers, though subscription prices remain unchanged.
Comments
Loading...