TTS is two studios

Google did not ship one text-to-speech model on September 23, 2026. It shipped two studios that share a schema and fight for the same budget line.

The Keyword post from Leland Rechis and Alan Cowen introduces Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS as the expressive generation pair in the Gemini Audio family. Model docs name the strings you actually call: gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. Flash is deep creative direction, character design, and long-form multi-speaker work across 130 languages. Flash-Lite is high-volume, cost-efficient dubbing, read-aloud, and agent cascades across 101 languages, and the recommended replacement for gemini-3.1-flash-tts-preview. Same speech API shape. Different workload. Marketers and creative ops who still buy “Google TTS” as one SKU will staff brand ambassador voice, scale localization, and agent cascades as if they were the same seat.

This is a primary-source review. No hands-on theater. Distinct from Vault’s Live dual-SKU note, the Codex /voice review, and the Claude default / classifier pieces. The question: which model string belongs on which lane this week.

What it is

Gemini TTS means text in, audio out, with fine-grained control over style and sound. Google’s speech-generation docs draw the line against the Live API: Live is interactive conversation; TTS is exact recitation for podcasts, audiobooks, dubbing, read-aloud, and scripted agent turns. Confusing those two is how a narration budget gets judged on interrupt latency it was never built for.

Both 3.8 TTS models accept text-only input and return audio-only output. Both support single- and multi-speaker generation, Voice design from prompts, Voice replication from reference audio, the Extended Voice Library (GET /v1beta/voices), and the same speech_metadata plus inline vocal-tag schema. Token limits on both cards: 8,192 input and 16,384 output. Batch, Flex, and Priority inference are supported on both. Availability on launch day (Keyword): Gemini API and Google AI Studio for developers; Flash-Lite enterprises coming soon via Gemini Enterprise API; Google Vids and Gemini Notebook named as product homes.

What changed

Two named generation SKUs, not one umbrella. The marketing title says “Gemini 3.8 text-to-speech.” The docs ship two endpoints. OrcaRouter’s September 23 note: there is no callable model named “Gemini 3.8 TTS.” Launch language only. The strings that resolve are Flash and Flash-Lite.

Flash as the creative studio. Keyword: create voices from scratch with natural-language prompts, direct line-by-line acting cues, long-form generation with minimal speaker drift, native two-speaker scene staging, scripted vocal bursts and backchanneling. Model docs: maximum fidelity, acting nuance, dialect coverage; best for audiobooks, studio narration, complex multi-speaker dialogue, heavy vocal-burst acting, difficult pronunciation, regional dialects. 130 languages.

Flash-Lite as the scale studio. Keyword: high-volume dubbing, audio content creation, expressive voice agents. Model docs: high throughput, low latency, cost efficiency; best for high-volume production, real-time voice agent cascades, read-aloud, voice replication, everyday single-speaker generation. 101 languages. Explicit recommended replacement for gemini-3.1-flash-tts-preview.

Voice surface expands. Keyword: generative voice design; 2,000+ production-ready voices including regional varieties (Mexican Spanish, Quebec French, Scots English); 30-second voice replication with consent verification, SynthID watermarking, and C2PA credentials; save/manage custom voices; voice remixing coming soon. Speech-generation docs: 30 prebuilt studio voices, Extended Voice Library filtering, persistent voice_... IDs (200 per project, 1-year TTL), optional stateless voicekey_... keys (7-day TTL).

Schema migration from 3.1 preview. Input text is a verbatim transcript. Stage directions like “Say cheerfully:” or “Speaker 1:” can be spoken aloud if left inline. Move sustained delivery into speech_metadata.style and speaker labels into speech_metadata.speaker. Keep angle-bracket tags for point-in-time vocal events only. Every multi-speaker turn must name its speaker. Unary requests now default to WAV (audio/wav) with a RIFF header instead of headerless PCM audio/l16; pipelines that wrapped PCM must stop double-wrapping, or set response_format for L16 / mu-law / A-law.

Benchmark claims. Keyword: Flash TTS #1 overall on Hume AI’s Voice Design Benchmark (71.4) and accent modeling (60.8); Flash and Flash-Lite #1 and #2 on Hume’s Overall Quality Index; strong Voice Arena preference in languages including Japanese, Brazilian Portuguese, Vietnamese, MSA, Mexican Spanish, and Hindi. SiliconANGLE (September 23) restates the Hume ranks without inventing rates.

Dual SKU table in prose

Lane A: gemini-3.8-flash-tts. Fidelity, acting, dialect coverage. 130 languages. Workloads: brand ambassador / character voice, long-form series, dual-speaker podcast or screenplay staging, minority dialects, heavy vocal bursts. Paid Standard through December 31, 2026: $0.50 per 1M text input, $9.00 per 1M audio output (about $0.00225 per 10 seconds). From January 1, 2027: $1.00 / $18.00. Batch and Flex cut audio in half; Priority raises it ($0.90 / $16.20 through year-end 2026).

Lane B: gemini-3.8-flash-lite-tts. Throughput, latency, cost. 101 languages. Workloads: bulk dubbing, read-aloud, high-volume agent cascades, everyday single-speaker, migration off gemini-3.1-flash-tts-preview. Paid Standard through December 31, 2026: same $0.50 text, $6.00 audio (about $0.0015 per 10 seconds). From January 1, 2027: $1.00 / $12.00. Input list matches Flash; the gap is the audio bill (about one-third cheaper on Standard through year-end).

Shared surface. Same schema, so switching is a model-string change. Both support Voice design, Voice replication, library voices, dual-speaker conversational mode, streaming, and caching. Multi-speaker in one request is capped at two speakers using prebuilt voices; custom designed or replicated voices in multi-character dialogue need per-speaker synth and 24 kHz PCM concat.

Macro cyan-lit microphone mesh grille on charcoal background

Consent, watermark, geo limits

Voice replication is not a free “clone whoever” button. Keyword: users must provide a verbal consent recording from the voice owner that matches the reference speaker before a replicated voice can be created. Every Gemini Audio clip is watermarked with SynthID. SiliconANGLE adds that Google attaches a C2PA record describing generation and modification metadata. Keyword: Voice replication through AI Studio is not available in Illinois, Texas, EEA, UK, Switzerland, and India. Creative ops with talent across those geos need a consent and residency map before promising replicas. Custom voice storage is a quota: 200 stateful voices per project, one-year TTL; stateless keys live seven days. Voice remixing remains “coming soon.”

What works / what breaks

What works. One schema, two dials: prototype on Flash, drop volume lanes to Flash-Lite without rebuilding request shape. OrcaRouter notes Artificial Analysis Voice Arena Elo intervals for the two overlap on their September 24 capture, which argues for listening tests over blind premium spend. Voice design, turn-level style, dual-speaker staging, and vocal-burst tags give producers a director-style surface. Flash-Lite is the named 3.1 preview replacement, and its $0.50 / $6.00 Standard card through year-end beats the still-listed 3.1 preview row ($1.00 / $20.00). Consent, SynthID, C2PA, and geo blocks are documented for a brand-safety appendix.

What breaks. Buying “Google TTS” as one line item overpays studio work on Lite or burns Flash rates on cascades. Language gaps on Lite are real: 101 versus 130. Check the tables before locking localization to Lite for cost. Schema footguns: inline stage directions get spoken; missing speaker fails multi-speaker turns; unary WAV default breaks code that always prepended a RIFF header to raw PCM. Custom voices in multi-character scenes need stitch steps (two-speaker single-request paths want prebuilt voices). AI Studio replication is blocked in listed regions; no consent clip means no replica. TTS is not Live: barge-in conversation belongs on Live SKUs.

Who it is for

Primary: Creative ops, localization leads, and growth marketers who own brand voice, dubbed video, podcast / audiobook production, or scripted agent speech and need to assign Flash versus Flash-Lite by workload.

Secondary: Platform engineers migrating off gemini-3.1-flash-tts-preview who need the schema checklist and rate cards.

Not the buyer: Anyone shopping one “Google voice” seat for both interruptible phone agents and 40,000-word audiobook fidelity.

Bright line vs Live SKUs

Do not rehash Vault’s Live dual-SKU piece. Hold the fence.

TTS (gemini-3.8-flash-tts, gemini-3.8-flash-lite-tts): text to speech, exact script, style metadata, studio and cascade generation. Priced on TTS text-in / audio-out rows.

Live (gemini-3.8-live, gemini-3.8-live-extended-thinking, related Live preview paths): real-time audio-to-audio conversation, multimodal inputs, agent loops. Separate Live pricing table.

Speech-generation docs: Live for dynamic conversation; TTS for exact recitation. Flash versus Flash-Lite is a generation-tier choice inside TTS. Live versus Live Extended Thinking is a conversation-mode choice inside Live. Mixing those four names into one “Gemini voice” budget is how you mis-staff both. Also not this review: Codex CLI /voice, Claude default flips, or classifier billing.

Risograph collage with blank ledger strip, magenta blot, and cyan circles on newsprint

Pricing and limits

Paid Standard through December 31, 2026 (pricing page, accessed September 24, 2026):

  • Flash TTS: $0.50 / 1M text in, $9.00 / 1M audio out (~$0.00225 per 10s). Caching: $0.125 input caching plus $0.50 / 1M tokens per hour storage through year-end.
  • Flash-Lite TTS: $0.50 / 1M text in, $6.00 / 1M audio out (~$0.0015 per 10s). Same caching shape.

January 1, 2027 step-ups double Standard text and audio on both ($1.00 / $18.00 Flash; $1.00 / $12.00 Lite). Batch and Flex are half Standard audio. Priority is higher ($16.20 Flash / $10.80 Lite audio through year-end 2026). Free tier lists free-of-charge input and output on these Standard rows.

Migration contrast only: Gemini 3.1 Flash TTS Preview remains at $1.00 text / $20.00 audio Standard (Batch $0.50 / $10.00). No invented rates. Both 3.8 cards list 8,192 / 16,384 token limits; Live API is not supported on TTS models.

What to do this week

  1. Name the two strings. Put gemini-3.8-flash-tts on brand ambassador, character, long-form, and hard-dialect lanes. Put gemini-3.8-flash-lite-tts on dubbing, read-aloud, and agent cascades. Kill the single “Google TTS” procurement label.
  1. If you are on gemini-3.1-flash-tts-preview, migrate schema before swapping the model. Move styles and speakers into speech_metadata, keep angle brackets for bursts only, stop double-wrapping WAV headers, land on Flash-Lite unless language or fidelity forces Flash.
  1. Run a same-script A/B on Flash versus Flash-Lite for your top three languages. Shared schema makes the test cheap. Prefer ears over Elo when the audio rate gap is ~33%.
  1. Map voice replication geos and consent. Illinois, Texas, EEA, UK, Switzerland, and India block AI Studio replication (Keyword). Build consent-clip capture into the talent workflow before promising replicas.
  1. Separate TTS generation budgets from Live conversation budgets. Do not judge Flash/Flash-Lite on barge-in latency, or Live on audiobook drift.
  1. Cap custom voice sprawl. 200 stateful voices per project and a one-year TTL mean brand libraries need an owner.
  1. Update the creative ops FAQ from primaries. Link the September 23 Keyword post, both model docs, speech-generation, and pricing TTS sections. One paragraph: two studios, same schema, 130 vs 101, Flash for depth / Lite for scale, consent + SynthID + C2PA, geo blocks, not Live.

Sharp close

TTS is two studios. Flash is the deep room where you design the ambassador and hold a long scene. Flash-Lite is the floor where you ship volume without lighting the audio meter on fire. Same API shape on purpose: mis-label the workload, pick the wrong string, and either overpay or under-act. This week, write the two model IDs into the runbook, keep Live off the same slide, and treat consent geos like production constraints, not footnotes.