Thursday, September 24, 2026
Follow on Google News

Google introduces two new text-to-speech Gemini models: Gemini 3.8 Flash TTS and 3.8 Flash Lite TTS

Google introduces two new text-to-speech Gemini models: Gemini 3.8 Flash TTS and 3.8 Flash Lite TTS

Yesterday, Google introduced two new text-to-speech models to the Gemini family: Gemini 3.8 Flash TTS and 3.8 Flash Lite TTS. These models transform voice generation from static presets into a dynamic creative studio.

Gemini 3.8 Flash TTS & Gemini 3.8 Flash Lite TTS

  • Gemini 3.8 Flash TTS: Built for deep creative direction and character design. Create entirely new voices from scratch using natural language prompts to bring characters to life across gaming, immersive audiobooks, podcasts, and interactive media. Direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling.
  • Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. Optimised for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance.

The Gemini 3.8 Flash TTS model powers a full vocal studio that enables users to create and use expressive, natural-sounding voices for every moment.

Features of these models

Create and customise your own voices

  • Generative voice design: With Gemini 3.8 Flash TTS, create bespoke voices from scratch by customising role, accent and voice characteristics across more than 100 languages and dialects using natural language prompting.
  • Expansive voice library: Access 2,000+ production-ready voices with broad language coverage — including regional varieties like Mexican Spanish, Quebec French, and Scots English.
  • Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.
  • Save and scale: Save and manage the custom voices you designed to ensure consistent performance and minimal drift across ongoing projects.
  • Voice remixing: Coming soon, pick a voice from our voice library and fine-tune timbre, pitch, pace, and accent. Use prompts to dial in characteristics (e.g. “add subtle Southern US accent” or “soften the delivery”).

Direct the performance, line by line

Both models give precise control over how each line is delivered.

  • Direct performance line by line: Write your own stage directions or let Gemini steer delivery with natural script cues.
  • Long-form generation: Maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift.
  • Native two-speaker scene staging: Direct multi-turn conversations seamlessly from a single script, while keeping both voices distinctly separated with natural conversational turn-taking.
  • Scripted vocal bursts & backchanneling: Add realistic conversational texture using nonverbal cues (like <laughs>, <sigh>, <gasp> and active-listening interjections (like |mhm| or|yeah|) for precise comedic timing and reaction beats.

Get expressive, high-quality speech generation built for global scale

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS enable truly expressive performances without sacrificing reliability, also securing the #1 and #2 spots respectively on Hume AI’s Overall Quality Index. The model shows major improvements on a wide range of use cases such as long-form content and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS.

In blind human preference evaluations on Voice Arena, Gemini 3.8 Flash and Flash-Lite TTS secure top positions amongst competitors in key global languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With support for over 100 languages, these models empower creators, developers, and enterprises to build high-quality, multilingual voice experiences worldwide.

Build with trust, consent, and transparency

For voice replication, Google’s system leverages consent verification: users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. Every audio clip generated by Gemini audio models is watermarked with SynthID.

Availability

Gemini 3.8 Flash TTS is rolling out:

  • For developers: In the Gemini API and Google AI Studio
  • For enterprises: Coming soon via API in Gemini Enterprise
  • For everyone: In Gemini Notebook

Gemini 3.8 Flash Lite TTS is rolling out:

  • For developers: In the Gemini API and Google AI Studio
  • For enterprises: Coming soon via API in Gemini Enterprise
  • For everyone: In Google vids
Add us as a preferred source on Google
Estuti Bajpai

Journalism student currently pursuing Masters in mass communication and Journalism from Chandigarh University. Well versed in public communication. Looking for an opportunity in the field of journalism to gain knowledge and use my journalism skills efficiently.

1 / 1