Google Gemini is pushing its text-to-speech technology beyond simply turning written words into a convincing voice.
The company has introduced two new speech-generation models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, with the two systems aimed at different kinds of audio production. Flash TTS is positioned for more creative, performance-driven work, while Flash-Lite TTS is designed for situations where businesses need to generate speech at a larger scale.
That distinction matters because the new release is not only about making synthetic speech sound better. Google is giving developers and production teams considerably more control over how that speech is performed, how many voices can appear in a conversation, and how those voices hold up during longer recordings.
More than 2,000 voices across 100-plus languages
One of the biggest changes is the size of the available voice library.
Instead of relying on a fixed catalogue of 30 legacy voices, developers can access more than 2,000 pre-built vocal profiles. The collection spans more than 100 languages and includes regional variations such as Quebec French, Scots English and Mexican Spanish.
That opens up considerably more room for applications where a generic English-speaking voice is not enough. Regional dubbing, multilingual customer interactions, translated content and locally produced audio can all benefit from having a wider choice of voices and linguistic variations.
Google is also preparing a voice-remixing capability that will allow audio teams to alter characteristics such as timbre, pitch, pace and accent contours using text instructions.
In practice, that moves text-to-speech closer to vocal direction. Instead of choosing a voice and accepting its default performance, creators will have more ways to describe how they want it to sound.
Scripts can control the performance, not just the words
Gemini 3.8 Flash TTS can also interpret performance cues written directly into a script.
Writers can include tags such as <laughs>, <sigh> and <gasp>, along with conversational interjections including |mhm| and |yeah|. The idea is to make generated dialogue less dependent on perfectly written sentences and more capable of reproducing the small sounds that naturally appear in human conversation.
For podcasts, games, audiobooks and character-driven content, those details could be particularly useful. A spoken scene often depends as much on pauses, reactions and delivery as it does on the actual words.
Two speakers can be handled inside the same script
Another notable capability is multi-speaker speech generation.
A single script can direct a conversation between two speakers while keeping their voices separate and maintaining conversational turn-taking over longer exchanges.
This is especially relevant for productions such as interview-style audio, fictional dialogue and conversational learning material, where manually generating each speaker separately can add extra editing work.
The models are also designed to maintain voice character and audio quality over extended recordings. According to the reported capabilities, the system can keep timbre stable across multi-hour audio, an area that matters for long-form formats such as audiobooks and episodic podcasts.
Gemini 3.8 Flash TTS and Flash-Lite serve different jobs
Google is not treating the two models as interchangeable.
Gemini 3.8 Flash TTS is geared towards use cases where creators want more direct control over a vocal performance. The examples include interactive entertainment, game development and long-form narration.
Gemini 3.8 Flash-Lite TTS, meanwhile, is aimed at higher-volume workloads such as automated media dubbing, conversational customer applications and translation pipelines.
The split suggests a practical trade-off between creative control and large-scale speech generation rather than a single model being expected to cover every production environment.
Early benchmarks put the models among the stronger TTS systems tested
The larger Gemini model also recorded strong results in the evaluations cited alongside the launch.
Gemini 3.8 Flash TTS received an overall score of 71.4 on the Hume AI Voice Design Benchmark, including a 60.8 score for accent modelling. In Hume AI’s Overall Quality Index, Flash TTS was placed first and Flash-Lite second in the evaluation.
The material also cites double-blind Voice Arena tests in which the models showed preference advantages across languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.
Those tests are useful indicators of comparative voice quality, although benchmark performance and listener preferences do not by themselves determine how a model will perform in every production setting.
Google is putting safeguards around custom voice creation
The ability to reproduce or customise voices inevitably raises questions around impersonation, and Google has built consent requirements into its voice-cloning process.
Creating a custom vocal profile requires a 30-second reference recording as well as a separate verbal consent recording from the owner of the voice. The two recordings are checked for acoustic alignment before the custom profile is processed.
Generated audio also carries SynthID audio watermarking and C2PA provenance metadata, giving detection systems a way to identify synthetic speech and trace its origin.
These measures are important because more convincing and more controllable synthetic voices also make responsible attribution increasingly necessary.
Where Gemini 3.8 Flash TTS is available
Developers can access both Gemini 3.8 Flash TTS and Flash-Lite TTS through Google AI Studio and the Gemini API. The models can also connect with developer frameworks including Agora, LiveKit, Pipecat and Vercel.
Commercial integrations named in connection with the rollout include Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang, with use cases spanning regional media translation and customer-service automation.
For end users, Flash TTS is being made available inside Gemini Notebook, while Google Vids uses Flash-Lite TTS. Administrative API access for Gemini Enterprise customers is expected in a later deployment phase.
Taken together, the release shows where Google’s voice technology is heading. Text-to-speech is becoming less about selecting a synthetic narrator and more about directing a performance. Larger voice libraries, regional language support, two-speaker conversations, long-form stability and written performance cues give creators more control over the final audio, while Flash-Lite provides a separate route for businesses that need speech generated at scale.
For developers, publishers and audio teams, that may be the more meaningful shift: the script is no longer limited to telling the model what to say. It can increasingly tell it how to say it.