Google has added computer-control capabilities directly to Gemini 3.5 Flash, allowing the model to see and operate computers, browsers, and mobile devices. The Decoder reports that the model reaches 78.4 on OSWorld, placing it near other frontier systems for computer-use tasks.
Developers can use the Gemini API to build agents for workflows such as software testing, browser automation, and office tasks. That moves computer use from a separate tool wrapper into the model’s native capabilities.
The release is another sign that AI agents are shifting from text-only assistants toward systems that can act in real user interfaces.