DeepSeek Ships V4-Pro-0813, Open-Sources Agent Harness, Raises API Prices Up to 52%
DeepSeek released build V4-Pro-0813 with major agent benchmark gains, open-sourced its Deepseek Harness agent framework under MIT license, and announced API price increases of up to 52% effective August 16.
DeepSeek has moved its flagship model out of preview with build V4-Pro-0813, simultaneously open-sourcing its agent software and raising API prices by as much as 52%.
The deepseek-v4-pro endpoint now runs the new build. Model name, parameter count, and the 1-million-token context window are unchanged, and DeepSeek says existing integrations require no changes. The model is available in DeepSeek's app and web interface under "Expert Mode," and now natively supports OpenAI's Responses API with Codex integration. Reasoning effort can be set to "low," "high," or "max," with DeepSeek recommending "high" for everyday agent use.
According to DeepSeek's own benchmark comparison, Terminal Bench 2.1 scores rose from 72.1 to 87.9, and DeepSWE scores climbed from 12.8 to 62.7. DeepSeek claims the new build beats Claude Opus 4.8 on several agent benchmarks, though it only outperforms Kimi K3 and Fable 5 in select categories.
Third-party benchmarking firm Artificial Analysis confirms the improvement but places it in context: V4-Pro rises from 45 to 53 on the Intelligence Index, tying GLM-5.2. That score still trails Muse Spark (57), Qwen 3.8 Max (58), Kimi K3 (60), and Claude Opus 5, which leads at 63. DeepSeek has not yet published weights for the new build; the April preview remains the only version on Hugging Face.
The update also responds to competitive pressure from within DeepSeek's own lineup. The smaller V4 Flash model, updated in late July with build 0731, nearly matched the Pro Preview's Intelligence Index score at a fraction of the cost.
Agent harness goes open source
DeepSeek Harness v0.1 ships as a Developer Preview under the MIT license, positioned as an open alternative to OpenAI's Codex and Claude Code. It runs on DeepSeek's new Cordis plugin system, where tools, sandboxes, sessions, and the UI are all swappable components. A continuous session log records every prompt, tool call, and result, allowing runs to be resumed, branched, and replayed. The project is led by Cui Tianyi, who joined DeepSeek from quantitative trading firm Jane Street in March 2026. A beta tester call in early August drew 712 project sign-ups within three days.
Prices rise, cache hits hit hardest
New rates take effect August 16 at 4:00 p.m. UTC. DeepSeek introduced peak/off-peak pricing tied to Chinese business hours (1–4 a.m. and 6–10 a.m. UTC), with off-peak rates costing half as much.
Off-peak, V4-Pro input rises from $0.435 to $0.66 per million tokens, and output climbs from $0.87 to $1.98. During peak hours, both rates double to $1.32 and $3.96. Cache hits see the steepest jump — from $0.003625 to $0.022 off-peak and $0.044 at peak — shrinking the cache discount from roughly 1/120th to 1/30th of the standard input rate. This change hits agents that repeatedly re-read the same files hardest, and reverses much of the price cut DeepSeek issued in May.
The price increase comes as DeepSeek raises new capital and prepares for an initial public offering.
What this means
DeepSeek is charging more for a model that, by independent measurement, still ranks behind four competitors on general intelligence. The bigger story may be the harness: giving away MIT-licensed agent infrastructure lowers the barrier for developers to build on DeepSeek's ecosystem even as usage costs rise, a bet that infrastructure lock-in matters more than raw benchmark leadership heading into an IPO.
Related Articles
OpenRouter Adds 'DeepSeek Pro Latest' Alias With 1M-Token Context Window
OpenRouter has introduced DeepSeek: DeepSeek Pro Latest, a routing alias that automatically points to whichever DeepSeek Pro model is newest. The endpoint offers a 1,049K token context window at $0.58 per 1M input tokens and $1.74 per 1M output tokens.
OpenRouter Adds DeepSeek Flash Latest Alias With 1M-Token Context Window
OpenRouter has launched deepseek-flash-latest, a persistent endpoint that always points to the current DeepSeek Flash model. It offers a 1,049K token context window, text-and-image input, and pricing of $0.15 per 1M input tokens and $0.60 per 1M output tokens.
DeepSeek Ships V4.1-Flash With Novel Encoder-Decoder Architecture, Cuts KV Cache to 1/8 of Predecessor
DeepSeek released V4.1-Flash, a 763B-parameter model built on a new causal encoder-decoder architecture that splits 8B active parameters for prefill and 16B for decode. The model adds native vision support, a 1M-token context window, and shrinks KV cache footprint to roughly 1/8 of DeepSeek V4 Flash, while retiring V4 Pro.
Vercel AI SDK Patch Adds Support for 'None' Reasoning Effort on GPT-6 Sol and Luna
Vercel released version 4.0.78 of @ai-sdk/openai, a patch that adds support for setting reasoningEffort to 'none' for GPT-6 Sol and Luna models. The update also introduces validation logic that warns on unsupported request-level effort updates and rejects invalid historical ones.
Comments
Loading...