Computer use in Gemini 3.5 Flash
Google has introduced computer use capabilities in Gemini 3.5 Flash, a new experimental feature that allows the AI model to perceive and interact with screen elements—such as clicking buttons, filling forms, and navigating interfaces—directly within a web browser, enabling automation of complex digital tasks without the need for custom integrations or APIs.
Background
- Google has released a feature called "computer use" in its Gemini 3.5 Flash AI model. In practical terms, this means the AI can directly control a computer interface — moving a mouse cursor, clicking on-screen elements, typing text, and navigating through software — rather than only processing text that a human pastes in.
- This type of capability (often called "GUI grounding" or "agentic" AI) is part of a broader industry push: competitors like Anthropic (with its "Computer Use" feature in Claude) and Microsoft (Copilot) are building similar models that can act as software agents, not just chatbots.
- For many technical and business users, "computer use" is significant because it suggests a path toward AI that can automate complex multi-step workflows — filling out forms, operating business software, or testing applications — without needing custom API integrations or direct database access.
- Google's 3.5 Flash is a lighter, faster variant of the Gemini model family, designed to balance cost and speed. Adding computer-use capability to this tier makes the feature more accessible for developers and everyday users compared to running a more expensive top-tier model.