Google has introduced Gemini 3.5 Transcribe, a speech-to-text model that can clean up dictation as it works. It automatically formats text, removes filler words such as “um,” and accepts custom vocabulary so unusual names or technical terms do not require repeated manual correction.

The model recognizes more than 85 languages, according to Google. For prerecorded audio, it can identify as many as three speakers and attach timestamps at the word level. Google positions it as a successor to Chirp 3, claiming improvements in multilingual transcription and word-error rates.

Two related audio models are also arriving. Gemini 3.5 Live is designed to handle interruptions, language detection and live visual input more reliably. An experimental Live version narrates progress while reasoning through more involved tasks.

The first consumer rollout is in English for the Gemini app on macOS and for the Rambler dictation feature on Android in selected countries and languages. Developers can try Transcribe in public preview through the Gemini API in AI Studio and Antigravity. Chrome support is planned but is not yet available.