News in Short
- Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.
- Flash TTS can create custom voices from natural-language prompts.
- Users can control delivery, pacing, accents and acting cues line by line.
- The models support more than 100 languages and dialects.
- Flash TTS offers voice replication from a 30-second audio sample with consent checks.
Google has introduced two new text-to-speech models designed to make AI-generated voices more expressive. Gemini 3.8 Flash TTS focuses on creative voice design, while Gemini 3.8 Flash-Lite TTS targets high-volume audio generation.
The models move beyond selecting a preset voice and generating speech. Google says creators can design voices, direct performances and build complete dialogue scenes using natural-language instructions.
Create Completely New AI Voices
One of the biggest capabilities is generative voice design. Gemini 3.8 Flash TTS can create a voice from scratch based on a natural-language description.
Creators can customise characteristics such as role, accent and voice style across more than 100 languages and dialects. Google gives examples ranging from dramatic fantasy characters to narrators with specific regional accents.
The model also includes access to more than 2,000 production-ready voices. These cover a wide range of languages and regional varieties, including Mexican Spanish, Quebec French and Scots English.
Replicate a Voice From 30 Seconds
Gemini 3.8 Flash TTS can also replicate a voice using a 30-second audio sample. Google says the feature can create a consistent vocal profile for a user’s own voice or a voice they have the rights to use.
Google has added safeguards around this capability. Voice replication requires consent verification from the voice owner, while generated audio includes SynthID watermarking and C2PA credentials.
Creators can also save custom voices for continued use. This is designed to maintain consistency across longer projects and reduce voice drift.
Direct How Every Line Sounds
The new models give creators control over individual lines instead of generating speech with a fixed delivery. Users can provide stage directions or natural-language cues to influence how the voice performs.
This can include pacing, acting style, pauses and other expressive elements. Google says creators can use the controls for scenarios ranging from customer service conversations to whispered suspense scenes.
The models also support scripted vocal effects and conversational reactions. Creators can add cues such as laughs, sighs and gasps, along with active-listening responses such as “mhm” or “yeah”.
Generate Full Conversations and Long Audio
Gemini 3.8 Flash TTS is not limited to short voice clips. Google says it can maintain voice quality, natural pacing and character timbre across hours of continuous audio with minimal speaker drift.
The system can also generate two-speaker scenes from a single script. It keeps the voices distinct while handling conversational turn-taking, which could be useful for podcasts, audiobooks and scripted content.
This gives creators a way to move from a written script to a more complete audio performance without separately producing every line.
Flash-Lite Targets High-Volume Audio
Google is also launching Gemini 3.8 Flash-Lite TTS for users who need audio generation at scale. The model is designed for high-volume dubbing, audio content creation and expressive voice agents.
It still provides control over tone, pacing and expressive nuance. However, its focus is on efficient generation across large volumes of audio rather than the deeper character-design capabilities of Flash TTS.
Google says the models have also performed strongly in its reported evaluations. Gemini 3.8 Flash TTS secured the top overall position in Hume AI’s Voice Design Benchmark, with a score of 71.4, according to Google.
Where Can You Use Gemini 3.8 TTS?
Developers can start experimenting with the new speech generation capabilities through Google AI Studio. Google describes the tool as a voice-design workspace where users can create new vocal identities or replicate eligible voices.
The Gemini API also makes the models available for developers building voice interfaces. Google says platforms including Agora, LiveKit, Pipecat and Vercel are enabling developers to integrate Gemini TTS into their applications.
Google also says the models are being used by companies working on dubbing, media localisation and conversational voice agents. For creators, the main change is the level of control. Instead of simply converting text into speech, Gemini 3.8 Flash TTS can turn a script into a directed audio performance.