DeepSeek V4-Pro got a new build on August 13, and its rate card changes on August 16 at 16:00 UTC. Output tokens go from $0.87 to $1.98 per million off-peak and $3.96 at peak, with some line items climbing more than 1,100%.
Key Takeaways
- The DeepSeek V4-Pro 0813 build moves from 45 to 53 on the Artificial Analysis Intelligence Index and still trails Claude Opus 5 at 63
- Billing splits into peak and off-peak windows on August 16 at 16:00 UTC, with cache hits taking the steepest increase
- DeepSeek open-sourced its agent software, Harness v0.1, under an MIT licence on the same day
A rebuilt model and a rebuilt rate card on the same day
DeepSeek published two things at once, and the second one takes back part of what the first one offers. The V4-Pro-0813 build is a clear step up, and the note shipped alongside it sets a price increase that lands three days later.
The gains hold up. Terminal Bench 2.1 rises from 72.1 to 87.9. DeepSWE jumps from 12.8 to 62.7, which reads less like progress and more like a model entering a category where it previously did not compete. The Artificial Analysis Intelligence Index moves from 45 to 53, level with GLM-5.2.
The ceiling is still visible. At 53 points, DeepSeek V4-Pro sits ten points below Claude Opus 5 at 63. Context stays at one million tokens, and the model picks up native support for OpenAI’s Responses API plus a Codex integration.
The increase was expected without being quantified. We flagged in early August that DeepSeek pricing was about to jump sharply, with no scale and no date attached at the time. Both blanks were filled at once in the general availability note published with the model.
The contrast with the spring positioning is sharp. Two weeks ago, DeepSeek V4 Flash was closing on GPT-5.6 Luna at 60% lower cost. That pricing argument has now been shelved across the whole line.
The honest reading of the DeepSWE jump fits in one sentence. A model going from 12.8 to 62.7 on an issue-resolution suite crosses into usable territory. The shift sits at the level of use case rather than degree, and that is exactly what makes the new rate card sting for teams that were about to adopt it.
What peak-hour billing forces engineering teams to do
The structural change is about price variability through the day more than the headline level. DeepSeek introduces two peak windows, 01:00 to 04:00 and 06:00 to 10:00 UTC, where the rate doubles against the rest of the clock.
Off-peak, input goes from $0.435 to $0.66 per million tokens and output from $0.87 to $1.98. At peak, input reaches $1.32 and output $3.96. The same request therefore costs up to four and a half times yesterday’s price depending on when it fires.
The rate card lands in a tight hardware context. We covered in July how DeepSeek is building its own AI chip to cut its Nvidia reliance. Charging double at peak hours amounts to managing a compute shortage through price.
The line moving most violently is the cache. A cache hit billed at $0.003625 per million becomes $0.022 off-peak and $0.044 at peak. The cache discount drops from one hundred and twentieth to one thirtieth of standard input pricing.
That detail is the real migration trap. Architectures that replay a long shared prefix on every call, typically agents carrying a heavy system prompt, were the ones getting the most out of that discount. They are the ones absorbing a twelvefold increase.
Two workarounds are already taking shape. Shifting batch jobs outside the two peak windows becomes worth engineering, since off-peak covers twenty hours out of twenty-four. Trimming the context re-sent on every turn becomes the other lever, now that the cache no longer absorbs it.
There is a second-order effect worth pricing in early. Peak hours cover the European morning, which is when most of the continent’s engineering teams push their heaviest interactive workloads. A rate card written around UTC therefore hits European users during their working day and rewards whoever can queue the job for later.
Teams that went through the July migration know the tempo. DeepSeek already imposed a tight schedule when V4 retired the old model names on July 24. Three days of notice for a full pricing overhaul fits the same operating style.
Alongside all this, DeepSeek opened Harness v0.1, its agent software, under an MIT licence in developer preview. It runs on the Cordis plugin system with swappable components. The open-source move arrives at the exact moment the API gets more expensive, which reads as something other than calendar luck.
More articles on Horizon
- Gemini 3.7 Flash Replaces 3.6 After Three Weeks
- Kimi K3 Test Ranks It First on Frontend Code
- Twitch Trains Amazon AI on Your Streams by Default
Chinese low cost stops subsidising the market
For two years the Chinese LLM play fit on one line: capability close to the American labs, at a fraction of the price. This announcement closes the second half of that offer.
The likely cause is industrial before it is commercial. Serving a model this size burns compute, and hourly pricing looks first like a load-shaping tool on saturated GPUs. The rate becomes a way to move demand around the clock rather than a plain margin adjustment.
For Chinese rivals, the window is open. Kimi K3 already closes the gap with top US models, and the cheap-price argument just changed hands without any of them having to move.
The open-sourced harness fits that reading too. Giving away the agent layer while charging more for the inference underneath is a familiar shape: the free component drives usage toward the metered one. DeepSeek is moving value from the software to the tokens, and the MIT licence is the cheapest way to make that trade look generous.
On the American side the effect is subtler. Downward pressure on pricing came largely from Chinese rates sitting in every engineering cost sheet as the reference point. When that reference climbs, US labs get more room on their own rate cards.
One deeper question stays open in the pricing note. If the Chinese entry rate was never sustainable at scale, then the real comparison between labs was never about the posted price these past two years. It was about the ability to hold that price, and that ability has just shown its limit.
Follow the story on Horizon.


