DeepSeek shipped a new build of its budget model on July 31, DeepSeek V4 Flash “0731”, and it now scores 50 on the Artificial Analysis Intelligence Index, one point behind GPT-5.6 Luna’s 51. The weights ship open under an MIT license, and a task runs roughly 60% cheaper than on OpenAI’s entry model. The capability-per-dollar line just moved again.
Key Takeaways
- DeepSeek V4 Flash “0731” climbs from 40 to 50 on the Intelligence Index, against 51 for GPT-5.6 Luna.
- Open weights under MIT, 284 billion total parameters with 13 billion active, a one-million-token context window.
- A task costs about 60% less than on GPT-5.6 Luna, with a cache discount pushed to 98%.
Have an AI Sum Up This Article
ChatGPTAn entry model catching the top tier
This is not a new architecture, it is a retrain. DeepSeek took its existing V4 Flash and re-post-trained it, keeping the 284 billion total parameters and 13 billion active. The weights published on Hugging Face under an MIT license confirm a one-million-token context window, unchanged.
The jump shows up in the numbers. On the Artificial Analysis Intelligence Index, the aggregate score most people use to rank models, the model moves from 40 to 50. GPT-5.6 Luna, the cheapest model in OpenAI’s lineup, tops out at 51. The gap between a Chinese open weight and the US leader’s entry model narrows to a single point.
The sharpest gain lands on the agentic side. On GDPval, the benchmark that measures how well a model chains real professional tasks, the score climbs from 1,189 to 1,559 Elo points. DeepSeek also flags a lower hallucination rate and consistent gains across every tested category, not one spike on a single metric.
For teams building agents, that detail matters more than the headline score. A cheap model that falls apart on multi-step work stays useless in production. A cheap model that holds up on GDPval becomes a real candidate to replace a proprietary model three times the price.
Price is the actual weapon
The real lever is not the benchmark, it is the bill. DeepSeek prices a task at roughly 60% less than the same task on GPT-5.6 Luna, and that is after OpenAI’s own 80% cut. The GPT-5.6 lineup had already slashed its entry price to 20 cents per million tokens, and DeepSeek still lands under it.
A second lever compounds the effect. The cache discount climbs to 98%, where the industry standard sits closer to 90%. For any workload that replays the same context often, a document-grounded assistant, an agent rereading the same files, the cached share of tokens turns almost free.
The new build also burns 12% fewer tokens than the previous one for an equivalent result. Fewer tokens emitted, a near-free cache, an already rock-bottom unit price: the three stack. At industrial volume, the cost gap with a proprietary model stops being an accounting line and becomes an architecture argument.
The shift extends a clear trend. Anthropic played the same card when Claude Opus 5 came in to match the top of the market at half the cost. The difference this time is that the price-breaker comes from a model whose weights anyone can download.
More articles on Horizon
- GPT-5.6 Luna: OpenAI Slashes the Price by 80%
- Gemini Robotics ER 2 Makes Robots Work Together
- Claude Mythos 5: Anthropic Admits Three Real Breaches
What MIT open weights force on rivals
The MIT license is the point that moves the debate. It allows commercial use with no strings, self-hosting included. A company that refuses to send its data to a third-party API can now run DeepSeek V4 Flash on its own hardware, at a quality close to OpenAI’s entry model.
The calendar does not help Western labs. The open-weight pressure out of China keeps building, as it did when Kimi K3 shipped 1.4TB of free weights. Each release of this kind chips away a little more at the exclusive-performance argument on the proprietary side.
The question now reaches straight into the majors’ strategy. Whether to open or lock a model is a front-line decision, to the point that Anthropic had to spell out its open-weights position in public. Holding a closed model gets harder to justify on technical lead alone.
On the competitive side, the expected response is mechanical. OpenAI already cut, Anthropic did too, and an open weight trailing Luna at 60% less reopens the round. The next lab to ship an entry model has to price it not against a proprietary rival, but against a free download.
The usual caveat stands. A one-point gap on an aggregate index says nothing about the specific cases where Luna keeps the edge, or about real inference speed on given hardware. But for a decision-maker weighing an API budget, the trajectory reads clean: quality is dropping toward free faster than closed-model prices are falling.
Follow the story on Horizon.


