Ordelvio News

AI · Published

Google Launches Gemini 3.8 Text-to-Speech Models

Google introduced Gemini 3.8 Flash TTS and Flash-Lite TTS, its most expressive models for custom voice creation and performance direction across more than 100 languages.

Isometric illustration of assembled objects on the theme of ai, gemini and text-to-speech, in electric blue, mint and pink.AI illustration
AI illustration for Ordelvio News (xAI grok-imagine-image-2.0) · not a photograph of the event · How AI is used

Google has introduced two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, which it describes as its most expressive audio generation models yet.

In a post by Leland Rechis, group product manager, and Alan Cowen, director of research science on behalf of the Gemini Audio Team, the company said the models are designed for creators, developers and enterprises. Gemini 3.8 Flash TTS supports deep creative direction and character design, while Gemini 3.8 Flash-Lite TTS targets high-volume, cost-efficient applications such as dubbing and voice agents.

The models allow users to create and customize voices from natural language prompts, scaling from an initial set of 30 voices to an infinite library. This includes generative voice design for bespoke characters across more than 100 languages and dialects by specifying role, accent and characteristics. An expansive library offers more than 2,000 production-ready voices with regional varieties. Voice replication from a 30-second audio sample is supported with consent verification, SynthID watermarking and C2PA credentials. Custom voices can be saved for consistency, with voice remixing capabilities coming soon.

Both models provide line-by-line performance direction using stage directions or natural script cues. They support long-form generation across hours of audio with minimal speaker drift, native two-speaker scene staging for conversations, scripted vocal bursts, backchanneling and non-verbal cues such as laughs, sighs and interjections.

Gemini 3.8 Flash TTS secured the top position on Hume AI’s Voice Design Benchmark with a score of 71.4 and leads in accent modeling with 60.8. The two models took the first and second spots on Hume AI’s Overall Quality Index. They also rank at the top in blind human preference evaluations on Voice Arena for languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.

The capabilities include safeguards such as consent verification for voice replication. Every generated audio clip contains an imperceptible SynthID watermark. The models complement the existing Gemini Audio family.

Rollout begins today in the Gemini API and Google AI Studio, with availability in Gemini Notebook and Google Vids. Enterprise access is coming soon via the Gemini API in Gemini Enterprise. Google is partnering with companies including Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang for integration into dubbing, localization and conversational agents.

  • ai
  • gemini
  • text-to-speech
  • voice-generation
  • google
  • audio-models