Gemini 3.8 Flash shipped on Wednesday and scores 73.7% on DeepSWE v1.1, three tenths of a point behind Claude Opus 5. Input stays at $0.75 per million tokens until December 31, yet the real cost per task climbs about 40% over the previous build. Google shipped a 3.8 Flash Cyber variant alongside it, handed only to defenders admitted into its new Fairwind programme.
Key Takeaways
- DeepSWE v1.1 moves from 65.3% to 73.7%, against 74.0% for Claude Opus 5
- $0.75 and $3.75 per million tokens until December 31, then double
- The Cyber variant hits 86.2% on CyberGym, ahead of GPT-5.6 Sol at 83.6%
Have an AI Sum Up This Article
ChatGPTThree Tenths of a Point Behind the Most Expensive Model on the Market
Google released Gemini 3.8 Flash and its 3.8 Flash Cyber sibling on Wednesday. The lab frames the first as its most intelligent workhorse, built for long-horizon software engineering, autonomous agents and multi-step enterprise workflows.
The benchmark table in Google’s official announcement of both models puts the gain on engineering work that runs long, not on isolated questions.
DeepSWE v1.1 climbs from 65.3% to 73.7% between 3.7 Flash and 3.8 Flash. Claude Opus 5 still leads at 74.0%, and GPT-5.6 Sol sits behind at 72.7%.
That three tenths of a point is the real signal here. A model billed at $0.75 per million input tokens now sits within reach of one billed at $5.00.
On the Artificial Analysis intelligence index, Gemini 3.8 Flash lands at 59 with high reasoning, matching GPT-5.6 Sol in its most demanding setting. HLE-Verified comes in at 54.9%.
For a team shipping code, that level moves the question. The choice no longer turns on raw capability but on what a single benchmark point justifies in extra spend.
The curve reads better next to the 3.7 Flash build that replaced 3.6 in mid-August. Every iteration picks up a few points while the listed price refuses to move.
Distribution landed immediately. The model already runs in AI Studio, Android Studio, the Gemini API, Gemini Enterprise, the Gemini app for Pro and Ultra subscribers, Google Search AI Mode and Sheets.
The Listed Price Hides What a Task Now Costs
Introductory pricing sits exactly where the previous generation left it, in line with the 3.6 Flash build that arrived cheaper than its predecessor in July. Input runs at $0.75 per million tokens and output at $3.75.
That grid carries an expiry date. It runs out on December 31, 2026, after which rates move to $1.50 in and $7.50 out, precisely double.
The number that matters to a product team sits elsewhere. At identical per-token pricing, average cost per task rises roughly 40%, from $0.40 to $0.58 against 3.7 Flash.
Google puts that gap down to added reasoning steps and iterative tool use. The model thinks longer and calls more tools, so it burns more for the same prompt.
Anyone who built a budget on per-token rates feels that directly. Migrating from 3.7 Flash at constant load costs more than before, with nothing changed on the pricing page.
The maths still works on long tasks. A model that clears a ticket in one pass where the last one needed three stays cheaper per result, heavier per-call invoice included.
Anthropic pulled the same lever days earlier, when Claude Fable 5.1 trimmed the bill by 25% without touching per-token pricing. Both labs are moving the discount off the published grid.
Something is missing from this calendar. Gemini 3.5 Pro and Gemini 4 are still absent, and a third Flash model in six weeks holds the ground instead, after a Gemini 3.5 Pro pushed back for a rebuild.
More articles on Horizon
- Claude Text Detection Opens to Media and Regulators
- Claude Fable 5.1 Codes Better and Costs 25% Less
- AI Video Has Replaced China’s Short Drama Actors
A Cyber Build Gated Behind Vetted Defenders
The second release of the day is narrower and more interesting. Gemini 3.8 Flash Cyber specialises in vulnerability discovery and automated patch generation.
On CyberGym it reaches 86.2%, past GPT-5.6 Sol at 83.6%. On real-world vulnerabilities, Google claims better than 70% success across twenty programming languages.
The patching side matters as much as detection. CWE-Bench shows 47.2% pass@1, and on Chrome Security the model produces 2.6 times more correct patches than leading commercial models.
Those numbers describe a capability, not a practice. Finding a flaw in a test repository and patching one in a live production base remain two very different exercises.
Google also flags robustness gains against prompt injection, measured on Gray Swan. A model reading hostile code all day has to resist instructions buried inside that code.
Access is locked down. The Cyber build ships through Fairwind, a programme reserving priority access for government authorities, critical infrastructure operators and software maintainers.
That gate sets a de facto norm. A model that finds flaws is no longer sold as a product, it is issued as a clearance, which the GPT-5.6 cyber build handed to defenders only had already sketched out.
Rivals now face pressure on two fronts at once. Matching Google means holding the line on long-horizon engineering while standing up a gated cyber track of their own.
One question stays open. A release cadence this tight on the budget tier keeps the doubt alive about the high end, and Google has given no date.
Follow the story on Horizon.


