Google has released two Gemini 3.8 text-to-speech models that turn written direction into custom voices and performed dialogue. Flash TTS is aimed at detailed character design, while the lower-cost Flash-Lite version targets high-volume dubbing, audio production and voice agents.

Flash can create voices through prompts covering role, accent and vocal traits across more than 100 languages and dialects. It can also reproduce a consistent voice from a 30-second sample when the user has permission. Both models accept line-by-line cues for pace, emotion, dialect shifts and conversational sounds, and can stage two speakers from one script. Google says long-form generation is designed to retain timbre across podcasts and audiobooks.

Voice replication requires a verbal consent recording that matches the reference speaker. Every generated clip receives Google’s imperceptible SynthID watermark and C2PA provenance credentials. Voice remixing, which will alter traits of library voices, is planned but not yet available.

Developers can use both models now through the Gemini API and Google AI Studio. Flash is also rolling out in Gemini Notebook, while Flash-Lite is coming to Google Vids. Enterprise API access through Gemini Enterprise is listed as coming soon.