DeepSeek to Quadruple API Prices for V4 Pro and V4 Flash Starting August 16
DeepSeek will raise API output token pricing roughly fourfold starting August 16, introducing peak and off-peak rates for its V4 Pro and V4 Flash models. Despite the increase, DeepSeek remains cheaper than competitors like OpenAI's GPT-5.6 Sol and Moonshot's Kimi K3.
DeepSeek is raising API pricing for its V4 model lineup by roughly four times starting August 16, ending the aggressive discount strategy that helped the Chinese AI lab undercut Western competitors.
According to the company, DeepSeek V4 Pro will cost $3.96 per 1 million output tokens during peak hours, up from the current $0.87 — a 4.5x increase. Off-peak pricing will run at $1.98 per 1 million output tokens, half the peak rate.
The smaller DeepSeek V4 Flash model will see output pricing rise from $0.28 to $1.32 per 1 million tokens at peak hours, with off-peak pricing set at $0.66.
DeepSeek said the new peak and off-peak structure is designed to "allocate resources more reasonably," suggesting the tiered pricing is meant to shift compute load away from high-demand hours rather than purely to increase revenue.
From promotion to permanent hike
The pricing shift reverses an earlier commitment. The discounted rates customers are currently paying were originally billed as a promotion set to expire May 31. In May, DeepSeek said it planned to make those discounted prices permanent. Instead, the company is now moving in the opposite direction, tying the increase to the rollout of its latest model, DeepSeek V4 Pro.
Still cheaper than rivals
Even after the fourfold increase, DeepSeek's pricing remains well below several competitors on a per-token basis. Moonshot AI's Kimi K3 costs $15 per 1 million output tokens — roughly 3.8x DeepSeek V4 Pro's new peak rate. OpenAI's flagship model, GPT-5.6 Sol, costs $30 per 1 million output tokens, nearly 8x DeepSeek's peak pricing.
OpenAI's lower-cost option, GPT-5.6 Luna, priced at $1.20 per 1 million output tokens, actually undercuts DeepSeek V4 Flash's new peak rate of $1.32, though it remains more expensive than DeepSeek's off-peak Flash pricing of $0.66.
What this means
DeepSeek's original disruption of the AI market rested heavily on radically undercutting Western labs on price while offering comparable performance. This increase — while still leaving DeepSeek cheaper than GPT-5.6 Sol and Kimi K3 — signals that the company's cost structure or business priorities have shifted enough to abandon a pledge to keep discounted rates permanent.
The peak/off-peak split is notable beyond the price hike itself: it points to real compute capacity constraints, forcing DeepSeek to use pricing as a demand-shaping tool rather than relying purely on flat-rate simplicity. Enterprise customers running high-volume workloads may now need to actively schedule inference around off-peak windows to capture savings, adding operational complexity that wasn't previously necessary.
For the broader market, the move suggests the era of DeepSeek functioning as an unambiguous low-cost anchor may be ending, even if it still undercuts most frontier-model competitors by a wide margin.
Related Articles
DeepSeek Ships V4-Pro-0813, Open-Sources Agent Harness, Raises API Prices Up to 52%
DeepSeek released build V4-Pro-0813 with major agent benchmark gains, open-sourced its Deepseek Harness agent framework under MIT license, and announced API price increases of up to 52% effective August 16.
DeepSeek Releases V4 Pro 0813 With 1.05M Token Context Window, Priced at $0.43/M Input
DeepSeek has shipped the general availability release of DeepSeek V4 Pro, codenamed 0813, featuring a 1,049,000-token context window. The mixture-of-experts model is priced at $0.43 per million input tokens and $0.87 per million output tokens, and is live now on OpenRouter.
DeepSeek Releases DeepSeek-V4-Pro-0813, a 1.7T-Parameter Model with DSpark Speculative Decoding
DeepSeek has released DeepSeek-V4-Pro-0813, a 1.7-trillion-parameter model that supersedes the DeepSeek-V4-Pro preview. The model adds a DSpark speculative decoding module and posts measurable gains on agentic and coding benchmarks, according to DeepSeek's technical report.
Meta Ships Muse Spark 1.2 Coding Model and Muse Code Agent, Undercuts Rivals with $0.20 Output Pricing
Meta released Muse Spark 1.2, a coding-focused upgrade to Spark 1.1, alongside Muse Code, its first dedicated terminal coding agent. The cheapest pricing tier drops output tokens to $0.20 per million, but requires users to share their data for training.
Comments
Loading...