DeepSeek Releases V4-Flash-0731, a 284B-Parameter Model That Beats Its Own Larger Pro Variant on Agentic Benchmarks
DeepSeek has shipped the full release of DeepSeek-V4-Flash-0731, a 284B-parameter model that according to DeepSeek outperforms its own larger V4-Pro (Preview) on agentic and coding benchmarks. Unsloth has published quantized GGUF versions, with lossless 8-bit weights requiring 162GB of storage.
DeepSeek has released DeepSeek-V4-Flash-0731, the official production version of DeepSeek-V4-Flash, superseding the earlier preview build. Unsloth has published GGUF quantizations of the 284-billion-parameter model, making it runnable on local and self-hosted hardware.
What's new
DeepSeek-V4-Flash-0731 shares its architecture with DeepSeek-V4-Flash-DSpark, including an attached speculative decoding module for faster inference. According to DeepSeek, the model shows "substantially enhanced agentic capabilities" compared to the preview release.
The most notable claim from DeepSeek: V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) across every listed benchmark despite activating far fewer parameters. DeepSeek describes the model as "broadly competitive with the strongest proprietary models available."
Benchmark results (as reported by DeepSeek)
| Benchmark | V4-Flash-0731 | V4-Flash (Preview) | V4-Pro (Preview) | GLM-5.2 | Opus-4.8 |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 82.7 | 61.8 | 72.1 | 81.0 | 85.0 |
| NL2Repo | 54.2 | 39.4 | 38.5 | 48.9 | 69.7 |
| Cybergym | 76.7 | 38.7 | 52.7 | — | 83.1 |
| DeepSWE | 54.4 | 7.3 | 12.8 | 46.2 | 58.0 |
| Toolathlon-Verified | 70.3 | 49.7 | 55.9 | 59.9 | 76.2 |
| Agents' Last Exam | 25.2 | 15.8 | 16.5 | 23.8 | 25.7 |
| AutomationBench Public | 25.1 | 10.8 | 12.8 | 12.9 | 27.2 |
| DSBench-FullStack | 68.7 | 37.0 | 41.8 | 61.8 | 71.6 |
| DSBench-Hard | 59.6 | 25.8 | 31.1 | 54.5 | 71.7 |
These figures come from DeepSeek's own technical report and have not been independently verified. DSBench-FullStack and DSBench-Hard are internal DeepSeek test sets. Code-agent evaluations used the minimal mode of the not-yet-released DeepSeek Harness, at max reasoning effort, temperature 1.0, top_p 0.95.
Opus-4.8 leads on most benchmarks, though DeepSeek-V4-Flash-0731 closes the gap substantially compared to the preview version and edges out GLM-5.2 on several coding and automation tasks.
Availability and quantization
Unsloth has released the model in GGUF format using its Dynamic 2.0 quantization method, which the company claims delivers superior accuracy relative to other quantization approaches. Two sizes are currently available:
- UD-Q4_K_XL: 155GB
- UD-Q8_K_XL: 162GB (described by Unsloth as "full precision, lossless")
Unsloth notes smaller quantizations are still in development. The model can also be run through Unsloth Studio, which exposes toggles for "High" and "Max" thinking modes. The model is licensed under MIT, and no inference providers currently host it.
Pricing for API access has not been disclosed, as DeepSeek has not announced a hosted endpoint for this specific release.
What this means
DeepSeek continues its pattern of releasing large models as open weights under permissive licenses, immediately enabling third parties like Unsloth to produce runnable quantized versions rather than waiting for official hosted access. The claim that a "Flash" variant beats its own "Pro" sibling on agentic benchmarks — if accurate — suggests DeepSeek is prioritizing efficient activation patterns over raw parameter scaling for agentic and coding workloads, a trend also visible in Chinese lab releases from GLM and others. The 155–162GB storage footprint still puts this squarely in enterprise or serious hobbyist hardware territory, not consumer GPUs, so broader adoption will likely wait on the smaller quantizations Unsloth says are coming.
Related Articles
DeepSeek Releases V4-Flash-0731, a 304B-Parameter Model Claiming to Beat Its Own Pro Preview on Agentic Benchmarks
DeepSeek has released DeepSeek-V4-Flash-0731, a 304-billion-parameter model that supersedes its earlier preview version with what the company describes as substantially enhanced agentic capabilities. According to DeepSeek's technical report, the model outperforms the larger DeepSeek-V4-Pro (Preview) on several coding and agent benchmarks despite a far smaller activated parameter count.
Unsloth Releases GGUF Quantizations of Kimi K3, a 2.8T-Parameter Open-Weight MoE Model
Unsloth has released GGUF quantizations of Kimi K3, a 2.8-trillion-parameter open-weight Mixture-of-Experts model from Moonshot AI with a 1-million-token context window and native vision support. The largest lossless quantization (Q8) weighs in at 1.56TB.
DeepSeek V4 Flash 'O731' Nearly Matches GPT-5.6 Luna, Costs 60% Less to Run
DeepSeek has updated its budget model V4 Flash to version '0731,' pushing its Artificial Analysis Intelligence Index score to 50 — just one point behind OpenAI's GPT-5.6 Luna — while costing an estimated 60 percent less per task. The MIT-licensed model keeps its 284B-parameter architecture but shows major gains in agentic benchmarks and token efficiency.
DeepSeek Releases V4 Flash 0731: 1M-Token MoE Model at $0.14/M Input Tokens
DeepSeek has released V4 Flash 0731, a sparse mixture-of-experts model with 13B active parameters out of 284B total and a 1049K token context window. The model targets coding, reasoning, and agent workflows, priced at $0.14 per million input tokens and $0.28 per million output tokens via OpenRouter.
Comments
Loading...