changelogDeepSeek

DeepSeek to Quadruple API Prices for V4 Pro and V4 Flash Starting August 16

TL;DR

DeepSeek will raise API output token pricing roughly fourfold starting August 16, introducing peak and off-peak rates for its V4 Pro and V4 Flash models. Despite the increase, DeepSeek remains cheaper than competitors like OpenAI's GPT-5.6 Sol and Moonshot's Kimi K3.

2 min read
0

DeepSeek is raising API pricing for its V4 model lineup by roughly four times starting August 16, ending the aggressive discount strategy that helped the Chinese AI lab undercut Western competitors.

According to the company, DeepSeek V4 Pro will cost $3.96 per 1 million output tokens during peak hours, up from the current $0.87 — a 4.5x increase. Off-peak pricing will run at $1.98 per 1 million output tokens, half the peak rate.

The smaller DeepSeek V4 Flash model will see output pricing rise from $0.28 to $1.32 per 1 million tokens at peak hours, with off-peak pricing set at $0.66.

DeepSeek said the new peak and off-peak structure is designed to "allocate resources more reasonably," suggesting the tiered pricing is meant to shift compute load away from high-demand hours rather than purely to increase revenue.

From promotion to permanent hike

The pricing shift reverses an earlier commitment. The discounted rates customers are currently paying were originally billed as a promotion set to expire May 31. In May, DeepSeek said it planned to make those discounted prices permanent. Instead, the company is now moving in the opposite direction, tying the increase to the rollout of its latest model, DeepSeek V4 Pro.

Still cheaper than rivals

Even after the fourfold increase, DeepSeek's pricing remains well below several competitors on a per-token basis. Moonshot AI's Kimi K3 costs $15 per 1 million output tokens — roughly 3.8x DeepSeek V4 Pro's new peak rate. OpenAI's flagship model, GPT-5.6 Sol, costs $30 per 1 million output tokens, nearly 8x DeepSeek's peak pricing.

OpenAI's lower-cost option, GPT-5.6 Luna, priced at $1.20 per 1 million output tokens, actually undercuts DeepSeek V4 Flash's new peak rate of $1.32, though it remains more expensive than DeepSeek's off-peak Flash pricing of $0.66.

What this means

DeepSeek's original disruption of the AI market rested heavily on radically undercutting Western labs on price while offering comparable performance. This increase — while still leaving DeepSeek cheaper than GPT-5.6 Sol and Kimi K3 — signals that the company's cost structure or business priorities have shifted enough to abandon a pledge to keep discounted rates permanent.

The peak/off-peak split is notable beyond the price hike itself: it points to real compute capacity constraints, forcing DeepSeek to use pricing as a demand-shaping tool rather than relying purely on flat-rate simplicity. Enterprise customers running high-volume workloads may now need to actively schedule inference around off-peak windows to capture savings, adding operational complexity that wasn't previously necessary.

For the broader market, the move suggests the era of DeepSeek functioning as an unambiguous low-cost anchor may be ending, even if it still undercuts most frontier-model competitors by a wide margin.

Related Articles

changelog

OpenRouter Adds 'DeepSeek Pro Latest' Alias With 1M-Token Context Window

OpenRouter has introduced DeepSeek: DeepSeek Pro Latest, a routing alias that automatically points to whichever DeepSeek Pro model is newest. The endpoint offers a 1,049K token context window at $0.58 per 1M input tokens and $1.74 per 1M output tokens.

changelog

OpenRouter Adds DeepSeek Flash Latest Alias With 1M-Token Context Window

OpenRouter has launched deepseek-flash-latest, a persistent endpoint that always points to the current DeepSeek Flash model. It offers a 1,049K token context window, text-and-image input, and pricing of $0.15 per 1M input tokens and $0.60 per 1M output tokens.

model release

DeepSeek Ships V4.1-Flash With Novel Encoder-Decoder Architecture, Cuts KV Cache to 1/8 of Predecessor

DeepSeek released V4.1-Flash, a 763B-parameter model built on a new causal encoder-decoder architecture that splits 8B active parameters for prefill and 16B for decode. The model adds native vision support, a 1M-token context window, and shrinks KV cache footprint to roughly 1/8 of DeepSeek V4 Flash, while retiring V4 Pro.

changelog

Vercel AI SDK Patch Adds Support for 'None' Reasoning Effort on GPT-6 Sol and Luna

Vercel released version 4.0.78 of @ai-sdk/openai, a patch that adds support for setting reasoningEffort to 'none' for GPT-6 Sol and Luna models. The update also introduces validation logic that warns on unsupported request-level effort updates and rejects invalid historical ones.

Comments

Loading...