VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Google ships Gemini 3.8 Flash TTS with 2,000-voice library and 30-second voice cloning

Google released Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, offering over 2,000 voices, 30-second voice cloning, and multi-speaker…

Google on September 23 released two text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which the company called its "most expressive audio generation models yet"1. The models support more than 100 languages and ship with a library of over 2,000 voices2.

Beyond the preset voice library, the models can create a custom voice from just a 30-second audio sample of a user's voice or a voice the user has the rights to use. The API also supports defining full conversations between multiple characters, each with different voices and voice-style instructions.

Simon Willison, who built a bring-your-own-key playground interface for the new models, reported that Gemini 3.8 Flash TTS generated 1 minute 18 seconds of audio in roughly 20 seconds at a cost of 2.74 cents. Willison noted that the underlying Gemini API has an open CORS policy, which enabled him to build the browser-based tool without a backend proxy.

ANALYSIS The combination of sub-three-cent generation costs, 20-second turnaround for over a minute of audio, and 30-second voice cloning lowers the practical barrier for developers integrating custom speech synthesis into applications.