Majormodel releaseGoogle DeepMind

Google Launches Gemini 3.8 Flash TTS with Custom Voice Creation Features

Published
Sep 23, 2026 — 17:39 UTC

Google Launches Gemini 3.8 Flash TTS

Google has introduced the Gemini 3.8 Flash TTS model, which allows users to create custom AI voices from text descriptions. The voice cloning feature can build a voice profile using just a 30-second audio sample. This model supports over 100 languages and offers more than 2,000 preset voices.

The pricing for audio output is set at $0.81 per hour for Flash TTS and $0.54 for Flash-Lite TTS through the end of 2026. Starting January 1, 2027, these prices will increase to $1.62 and $1.08 per hour, respectively. Additionally, both models charge $0.50 per million tokens for text input, while audio output costs $9.00 per million tokens for Flash TTS and $6.00 for Flash-Lite TTS.

Google claims that the generated audio clips carry an inaudible SynthID watermark to help detect AI-generated speech. The models are designed to produce hours of audio with minimal 'speaker drift,' and Google asserts that they lead in most categories of Hume AI's text-to-speech benchmark.

Developers using platforms like Agora, LiveKit, Pipecat, and Vercel can leverage these new capabilities to enhance their applications, particularly in creating personalized voice experiences. This follows Google's recent advancements in TTS technology, including the launch of Gemini 3.8 TTS with extensive voice options.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: The Decoder