Core Announcement and Key Timeline

On September 23, 2026, Google officially launched two new Text-to-Speech (TTS) models in the Gemini family: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Both models are available to developers immediately through the following channels:
- Developer access: Gemini API and Google AI Studio now host the new models
- Enterprise access: Gemini Enterprise API will follow shortly
- Consumer integration: Gemini Notebook includes Flash TTS; Google Vids includes Flash-Lite TTS
- Voice cloning limitations: Voice cloning via AI Studio is unavailable in Illinois, Texas, the European Economic Area, the UK, Switzerland, and India
From Static Voice to Creative Studio
The core advancement lies in transforming speech generation from selecting presets to programmable, expressive performance control. While Flash TTS targets deep creative scenarios (e.g., game roles, immersive audiobooks), Flash-Lite was optimized for high-volume, cost-efficient use cases like bulk dubbing and voice agents.
Both models enable line-by-line script control with identical foundational capabilities:
- Generative voice design: Create custom voices with natural language prompts across 100+ languages/dialects
- Expanded voice library: Access 2000+ ready-to-use voices, including regional variants like Mexican Spanish, Quebecois French, and Scottish English
- Voice cloning: Reconstruct voice features from 30 seconds of audio, protected by SynthID watermarking and C2PA credentials
- Scripted performance control: Adjust rhythm, emotion, and dialect per line via stage direction or Gemini prompts
A surprising data point: Flash-Lite, despite its cost-optimized positioning, achieved second place on Hume AI’s overall quality index, trailing only Flash TTS at first—proving performance reduction was carefully managed.
Technical Comparison and Benchmark Results
| Capability | Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS | Gemini 3.1 Flash TTS |
|---|---|---|---|
| Positioning | Deep creative/role design | High-volume/cost-efficient dubbing | Basic TTS |
| Voice customization | Create entirely new voices via prompts | Fine-tune pitch/rhythm/expression | Preset-only selection |
| Dual-speaker script control | Supported with significant improvements | Supported | Basic support |
| Hume AI overall quality index | 1st | 2nd | — |
| Hume AI accent modeling | 60.8 | — | — |
| Voice Arena multi-language blind test | Leading | Leading | — |
Gemini 3.8 Flash TTS ranked first in both total score (71.4 points) and accent modeling (60.8 points) on the Hume AI benchmark. In Voice Arena blind human preference tests across six major languages—Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi—the new models outperformed competitors in all categories.
Practical Adoption Guidance
- Early adopters should try: Podcast creators (need long-form continuity), game developers (dynamic multi-character speech), immersive content producers (real-time synthetic voice generation)
- Consider waiting if: You’re in restricted regions and need voice cloning; your enterprise requires ultra-low-latency batch generation until Gemini Enterprise launches
Final Thoughts
Speech synthesis is shifting from “functional” to “trustworthy” and “controllable.” By embedding SynthID watermarking and consent verification, Gemini 3.8 establishes a new equilibrium between creative freedom and identity protection—an evolution from tool to platform ecosystem.
