Gemini 3.8 Flash TTS
Free online text to speech that follows your direction. Gemini 3.8 Flash TTS voices scripts with emotion, dual speakers and 130 languages.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Gemini 3.8 Flash TTS: Flagship Speech Quality, Workhorse Economics
Launched September 23, 2026, the lineup pairs an expressive creative model with a low-cost tier tuned for high-volume audio jobs.
- A Single Release, Two Distinct JobsGemini 3.8 Flash TTS leads on expressiveness, while Flash-Lite TTS keeps per-minute costs down when you need speech at scale.
- Steer the Performance Instead of Loading a PresetPer-turn style prompts, structured speech metadata and inline vocal events control tone, tempo, feeling and regional accent.
- Invent a Voice or Clone One With PermissionWrite a voice into existence with plain words, or mirror a real person using a sample clip plus their matching consent audio.
How to Prompt Gemini 3.8 Flash TTS for Clean Results
Four habits that keep your transcript readable and let the delivery metadata do the acting.
What Gemini 3.8 Flash TTS Can Actually Do
From line-level acting control to 130 supported languages, here is a closer look at the flagship model's toolkit.
Line-Level Acting Control
Set style per turn and drop in laughs, sighs, coughs, breaths or pauses — it feels like coaching a performer rather than picking a stock voice.
Describe a Voice in Plain Words
Prompts can specify age range, personality, accent, timbre and role, backed by 2,000+ production voices available through the Voices endpoint.
Cloning Guarded by Consent
You need a clean reference sample plus matching consent audio from the same adult speaker, and output carries SynthID watermarking with C2PA credentials.
Built-In Two-Voice Scenes
Write the exchange once and the model handles podcasts, teaching dialogues, product walkthroughs and game scenes without stitching lines by hand.
Steady Quality Across Long Narrations
Google documents consistent identity, timbre, loudness and room tone through multi-minute narration and drawn-out conversations.
130 Languages With Local Accents
The flagship handles 130 languages versus Flash-Lite's 101, including regional accents, minority dialects and IPA pronunciation overrides.
Questions People Ask About Gemini 3.8 Flash TTS
Quick answers on cost per minute, model selection, benchmark scores and the safety rules behind voice replication.
What does Gemini 3.8 Flash TTS charge per minute?
Roughly 1.35 cents of audio per minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.
When should I pick Flash TTS over Flash-Lite TTS?
Choose Flash TTS when acting detail and long-form narration matter; pick Flash-Lite TTS for high-volume jobs and faster turnaround.
How does it stack up against rival voice models?
Google cites a 71.4 score on Hume's Voice Design Benchmark, and Voice Arena ranks it second with 1,260 Elo.
What changed from the Gemini 3.1 Flash TTS Preview?
Flash-Lite TTS takes over from the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.
Can I clone a voice, and what rules apply?
Yes, but only with a reference clip and a matching consent recording from the same adult speaker.
Why does the model read my stage directions aloud?
Because the input is treated as a verbatim transcript — shift lasting directions into the speech metadata fields instead.
Hear Gemini 3.8 Flash TTS Read Your Own Script
Try both tiers inside the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than a different model ID. Weigh batch against priority inference before locking in a production budget.
