Gemini 3.6 Flash Ships Cheaper as 3.5 Pro Slips

Gemini 3.6 Flash on a Google control desk facing a grounded 3.5 Pro spacecraft stuck on the floor

Google has rolled out three new Gemini Flash models led by Gemini 3.6 Flash, which trims token use by up to 17 percent and lands cheaper than the version it replaces. The flagship 3.5 Pro is still missing, and in the same window Google confirmed it has begun its most ambitious pretraining run yet for Gemini 4.

Key Takeaways

  • Gemini 3.6 Flash cuts token usage by up to 17 percent and undercuts its own predecessor on price.
  • A cheaper Flash-Lite and a governments-only Flash Cyber complete the release, while 3.5 Pro stays in testing.
  • Google has started pretraining Gemini 4, a sign the contest has moved onto cost and cadence rather than a single flagship.

Have an AI Sum Up This Article

ChatGPT

Three Flash models, and the Pro that never arrives

Google shipped the release as a trio. Gemini 3.6 Flash is the workhorse, tuned for coding, knowledge work and multimodal tasks, and it is the one most teams will actually wire into production.

Below it sits Gemini 3.5 Flash-Lite, framed as the most cost-effective option in the class. Above the everyday tier, Gemini 3.5 Flash Cyber is a variant fine-tuned to find and fix security vulnerabilities, and it ships on a tighter leash.

That leash is the detail worth flagging. Flash Cyber is restricted to governments and trusted partners under a limited-access pilot, which tells you Google reads security tuning as a capability it does not want loose on the open API just yet.

The gap in the lineup is the Pro. Gemini 3.5 Pro still has not shipped, and product lead Logan Kilpatrick noted the model is in testing with partners with a hope to land soon. The last update to that line dates back to February, which is a long time to leave the top slot open while the flagship keeps slipping and Google rebuilds it from a weaker base.

For a team choosing a model today the message is blunt. The frontier Gemini is not available, so the practical decision is between the Flash tiers, and that pushes the whole conversation off raw capability and onto running cost.


Gemini 3.6 Flash

Cutting tokens is how Google fights on price

The headline number is the 17 percent cut in token usage. A model that does the same job while emitting fewer tokens is directly cheaper to run, because inference is billed by the token, so the saving lands on every call a team makes.

For a product team building agents this is the number that matters more than a benchmark. An agent loops through a task, calling the model again and again, so a per-call discount compounds fast across a real workload and can move a use case from too expensive to viable.

Google was explicit about the target. It pitched the release around efficiency, latency and reliability for customers building AI agents at scale, which is a plain statement that the company is competing for the high-volume tier where most routine work now happens.

On the competitive side the timing is not gentle. Chinese labs have been dumping large models into subsidised preview, and we saw it again when Alibaba pushed Qwen 3.8 out at a tenth of the standard rate, planting a reference price the whole market now has to answer.

Google’s reply is to compete on efficiency rather than on a shock discount. Trimming tokens lowers the real cost without printing a promotional rate that has to be walked back later, which is a steadier way to hold the volume tier against pricing pressure from Hangzhou and Beijing.

The losers in that framing are the labs still charging premium rates for a general model, the bind Anthropic edged into when it kept Fable 5 free on Max while pushing Pro users onto per-token pricing. When the cheapest credible option keeps getting cheaper, every renewal conversation starts from a lower anchor, and the pressure lands hardest on whoever cannot match the cost curve.


More articles on Horizon


What the Gemini 4 run signals for the roadmap

The quiet headline is Gemini 4. Google confirmed it has kicked off its most ambitious pretraining run to date, which is a heavy commitment of compute and a clear statement about where the company is placing its next bet.

Read against the missing 3.5 Pro, the move looks deliberate. Polishing an interim flagship matters less if the real jump is a generation away, so shipping efficient Flash models now while the big pretraining run cooks is a coherent way to stay in the market without burning attention on a stopgap.

For rivals that sets an awkward clock. Anthropic and OpenAI now have to weigh their own release cadence against a Google that is holding the low-cost tier today and telegraphing a frontier model tomorrow, a squeeze that shows up in the same price war that pushed the largest open Chinese models to chase the top American systems on cost.

The open question is whether the cadence bet pays off. A pretraining run is a promise, not a product, and until Gemini 4 lands the frontier slot stays empty and the competitive read rests on models that already exist rather than one Google says is coming.

Usage over the next weeks will settle the price question first. Teams wiring Gemini 3.6 Flash into live agents will confirm whether the 17 percent figure holds on real workloads, and that is where the release either earns the volume tier or hands it back.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *