Gemini 3.5 Flash Sees Your Screen and Clicks for You

Gemini 3.5 Flash

Google has dropped native computer control inside Gemini 3.5 Flash, letting the model watch the screen and act on it directly. On the OSWorld benchmark, Flash now matches Sonnet 4.6, beats Gemini 3.1 Pro, and trails only Opus 4.8.

Key Takeaways

  • Gemini 3.5 Flash bakes Computer Use directly into the model itself.
  • OSWorld score of 78.4, tied with Sonnet 4.6 and above Gemini 3.1 Pro.
  • Two optional enterprise guardrails and adversarial training wrap sensitive actions.

Have an AI Sum Up This Article

ChatGPT

A Native Agent Wired Into the Model

Google is taking a clean step. Gemini 3.5 Flash now sees the screen, understands what is there and interacts with a computer on its own, without leaning on an external orchestrator. Computer Use lives inside the model.

The reach covers the three environments that matter. Web browsers, mobile apps and desktop machines all get the same treatment, with a perception-action loop the model drives end to end. That autonomy runs deeper still, Gemini 3.5 Flash powering autonomous agents.

The integration plugs into the rest of the Gemini toolbox. Function calls, Search and Maps stay available in the same request, letting the agent chain research, reasoning and execution without leaving its context. Google keeps stacking Gemini pieces, Nano Banana offering free images.

For first use cases, Google flags two practical fronts. Automated software testing and office automation open the playing field, two spaces already pursued by Codex over at OpenAI.

Access runs through the Gemini API and the Gemini Enterprise Agent Platform. To help developers kick the tires, Google ships a Browserbase demo and a reference implementation on GitHub, but commits to no public rollout timeline.


Gemini 3.5 Flash

78.4 on OSWorld: Flash Closes In on Opus 4.8

OSWorld is the arbiter. Gemini 3.5 Flash lands at 78.4, a clean jump from the 65.1 of Gemini 3 Flash. On that battlefield, the new version retires the old one without debate. The full picture is available in Google’s announcement introducing computer use in Gemini 3.5 Flash.

Against OpenAI, the result favors Google on the lightweight tier. GPT-5.4 mini sits at 72.1, while GPT-5.5 reaches 78.7. Flash therefore matches OpenAI’s flagship in this measurement.

Against Anthropic, the picture is more layered. Opus 4.8 still leads at 83.4 and keeps a clear margin. Sonnet 4.6 lands at 78.4, the exact score of Gemini 3.5 Flash, and Google now hits Sonnet-level numbers without the Sonnet bill.

The other detail that matters is Google’s internal stack. Gemini 3.1 Pro caps out at 76.2, below the new Flash. The premium model is now behind a Flash variant on agentic workloads, which reshuffles the pricing logic for developers.

On a productivity benchmark built around real actions, the chat-model hierarchy no longer holds. A Flash model armed with strong perception and tool use can stick to the top of the catalog, so long as the task lives inside the tested perimeter.


Also on Horizon:


Sandboxes and Safeguards: Google Steps Carefully

Agentic AI puts safety in the spotlight. Google trained Gemini 3.5 Flash with a dedicated adversarial layer and offers two optional enterprise guardrails on top.

The first guardrail enforces user confirmation on sensitive actions. The second turns on automatic detection of indirect prompt injection, a risk amplified by the fact that the model is reading whatever happens to be on the screen.

The integration guidance is explicit. Google recommends sandboxing, human oversight on critical workflows and strict access controls, in line with the posture already taken by OpenAI for its ChatGPT agent.

In the short term, adoption will hinge on integrators. The first targets are QA teams and modern RPA shops, who can industrialize an agent able to test a user journey or push a workflow inside a SaaS without a dedicated connector.

In the medium term, the mix of a fast Flash model and a mature agent platform changes the math on RPA and automation pricing. If the OSWorld trajectory holds in real-world tasks, Google takes the lead on the low-cost agent market, while Anthropic keeps aiming at the high end. Google spreads the agent everywhere, down to Remy, its personal AI agent.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *