Gemini 3.8 Flash TTS

Free online text to speech that follows your direction. Gemini 3.8 Flash TTS voices scripts with emotion, dual speakers and 130 languages.

Gemini 3.8 Flash TTS
Shape richly acted voice tracks, or keep costs low at scale with the lighter Flash-Lite tier
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Gemini 3.8 Flash TTS: Flagship Speech Quality, Workhorse Economics

Launched September 23, 2026, the lineup pairs an expressive creative model with a low-cost tier tuned for high-volume audio jobs.

  • A Single Release, Two Distinct Jobs
    Gemini 3.8 Flash TTS leads on expressiveness, while Flash-Lite TTS keeps per-minute costs down when you need speech at scale.
  • Steer the Performance Instead of Loading a Preset
    Per-turn style prompts, structured speech metadata and inline vocal events control tone, tempo, feeling and regional accent.
  • Invent a Voice or Clone One With Permission
    Write a voice into existence with plain words, or mirror a real person using a sample clip plus their matching consent audio.

How to Prompt Gemini 3.8 Flash TTS for Clean Results

Four habits that keep your transcript readable and let the delivery metadata do the acting.

What Gemini 3.8 Flash TTS Can Actually Do

From line-level acting control to 130 supported languages, here is a closer look at the flagship model's toolkit.

Line-Level Acting Control

Set style per turn and drop in laughs, sighs, coughs, breaths or pauses — it feels like coaching a performer rather than picking a stock voice.

Describe a Voice in Plain Words

Prompts can specify age range, personality, accent, timbre and role, backed by 2,000+ production voices available through the Voices endpoint.

Cloning Guarded by Consent

You need a clean reference sample plus matching consent audio from the same adult speaker, and output carries SynthID watermarking with C2PA credentials.

Built-In Two-Voice Scenes

Write the exchange once and the model handles podcasts, teaching dialogues, product walkthroughs and game scenes without stitching lines by hand.

Steady Quality Across Long Narrations

Google documents consistent identity, timbre, loudness and room tone through multi-minute narration and drawn-out conversations.

130 Languages With Local Accents

The flagship handles 130 languages versus Flash-Lite's 101, including regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Questions People Ask About Gemini 3.8 Flash TTS

Quick answers on cost per minute, model selection, benchmark scores and the safety rules behind voice replication.

1

What does Gemini 3.8 Flash TTS charge per minute?

Roughly 1.35 cents of audio per minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.

2

When should I pick Flash TTS over Flash-Lite TTS?

Choose Flash TTS when acting detail and long-form narration matter; pick Flash-Lite TTS for high-volume jobs and faster turnaround.

3

How does it stack up against rival voice models?

Google cites a 71.4 score on Hume's Voice Design Benchmark, and Voice Arena ranks it second with 1,260 Elo.

4

What changed from the Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS takes over from the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what rules apply?

Yes, but only with a reference clip and a matching consent recording from the same adult speaker.

6

Why does the model read my stage directions aloud?

Because the input is treated as a verbatim transcript — shift lasting directions into the speech metadata fields instead.

Hear Gemini 3.8 Flash TTS Read Your Own Script

Try both tiers inside the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than a different model ID. Weigh batch against priority inference before locking in a production budget.