DeepSeek has released V4-Flash-Vision-Exp, an experimental multimodal variant of its fast model that claims near parity with Opus 4.8 on agent tasks. The published tables read more carefully than the headline: three wins out of eleven benchmarks, and a gap that stays wide on repository-scale coding.
Key Takeaways
- V4-Flash-Vision-Exp bolts image understanding onto V4-Flash without touching its text ability or its agentic behaviour.
- It beats Opus 4.8 on three of eleven benchmarks, including Agents’ Last Exam and ZeroBench, and trails by twelve points on NL2Repo.
- Images are capped at 384 tokens each at V4-Flash rates, with a ceiling of 600 images per request.
Have an AI Sum Up This Article
ChatGPTA vision layer bolted onto V4-Flash, not a fresh foundation
What shipped on August 21 is an extension rather than a rebuild. V4-Flash-Vision-Exp takes the text model V4 Flash, the one that landed a point behind GPT-5.6 Luna at 60% lower cost, and teaches it to read images.
Reasoning, general knowledge and above all the agentic behaviour of the base model carry over untouched. That is the commercial pitch: same engine, new eyes.
The model is already served on the API platform under the identifier deepseek-v4-flash-vision-exp. Images go in as base64, as an external URL or through the Files API, in JPEG, PNG, GIF and WebP.
The integration choice says something about the strategy. Rather than spinning up a separate multimodal family with its own release cycle and its own price sheet, DeepSeek pushes vision into a tier customers have already wired in, where the integrations exist and the cost per call is known.
The Exp suffix is not decoration. This is an experimental milestone, with results measured inside the company’s own harness in minimal mode, and no independent verification available at this point.
The API documentation spells out the 600-image ceiling per request and the free Files API, two parameters that matter more than a benchmark point to anyone pushing document batches through a pipeline.
Three benchmark wins out of eleven, and a 384-token cap per image
On DeepSeek’s own table the model edges past Opus 4.8 on three lines. It takes DeepSWE by 1.3 points, Agents’ Last Exam by 1.6 with 27.3 against 25.7, and ZeroBench at Pass@5 by a single point, 35.0 against 34.0.
On the other eight lines it sits behind, sometimes barely. Terminal Bench 2.1 puts it at 83.9 where Opus 4.8 holds 85.0, and Chartography at 64.3 against 65.0.
Elsewhere the distance is real. On NL2Repo, which measures repository-scale code generation, DeepSeek scores 57.7 against 69.7. On DSBench-Hard the deficit runs to roughly eight points.
The genuine jump shows up against its own predecessor. On ApexBench at Pass@1 the vision variant reaches 36.5 where V4-Flash 0731 topped out at 26.2, which lands it three points off Opus 4.8 on that specific line.
On billing, every image is capped at 384 tokens, charged at the rates already running on V4-Flash. Costing out a document pipeline becomes predictable, which is not the norm across this market. Elsewhere in the lineup the adjustment runs the other way, since DeepSeek shifted V4-Pro to peak and off-peak billing with cache line items climbing more than 1,100%.
That cost discipline is a house habit by now. DeepSeek had already shipped DSpark, the speculative decoding framework that lifts its model speed 60 to 85% on GPU. The lab had also raised 7.4 billion dollars at a 50 billion valuation, with Tencent and CATL leading the round, which is what funds pricing like this.
More articles on Horizon
- Claude Security Now Scans Code With Mythos 5
- OpenAI Closes the Gap With Anthropic in Business
- AI Writing Shows Up on One in Three Web Pages
What it unlocks for agent teams, what it costs closed models
For a team running agents, the podium is not the point. The point is handling screenshots, scanned PDFs and charts inside the same loop as text, without routing out to a second, pricier model.
Keeping the base model’s agentic behaviour serves exactly that. Migrating from V4-Flash means no tool rewrite and no system prompt recalibration, which is worth more in practice than 1.6 points on a leaderboard.
The caveat sits in the experimental label and the in-house numbers. A score produced in your own harness is not a reproduced result, and DeepSeek flags that itself.
Competitively, the pressure lands on multimodal pricing rather than peak capability. Anthropic keeps its edge on repository-scale code, yet becomes harder to justify on visual agent work where the gap narrows to tenths of a point.
Worth asking who loses most here. Vendors selling a separate vision component, billed on top of a text model, watch that revenue line get squeezed by a model doing both at one rate with a hard cap per image.
Its own pricing does not move in one direction either. The company tunes margins tier by tier rather than across the catalogue, and the vision variant lands while it tests how high the bill can go elsewhere.
It has the balance sheet for it, and funding this kind of experiment no longer depends on closing a round. What it is short of is measured in compute capacity rather than cash.
What decides the next chapter is silicon. DeepSeek has spent a year building its own inference chip to loosen its grip on Nvidia, and serving a vision variant at this price point burns precisely the kind of capacity Chinese labs are short of today.
Follow the story on Horizon.


