ElevenLabs Launches Eleven v4 Speech Model with Enhanced Expressiveness
- Published
- Sep 29, 2026 — 14:45 UTC
On September 29, 2026, ElevenLabs released its Eleven v4 speech model, achieving a pronunciation benchmark score of 91.7%, up from 85.6% in v3. The model supports over 90 languages and can handle requests of up to 10,000 characters. The Turbo variant boasts a response time of 150 milliseconds, significantly outperforming competitors like Cartesia Sonic 3.6 at 262 milliseconds and OpenAI's GPT-4o mini TTS at 814 milliseconds.
In blind tests, 65% to 81% of listeners preferred the v4 model, which utilizes a new architecture that analyzes a script's tone, pacing, and context. ElevenLabs claims that cloned voices can now speak multiple languages with native accents, addressing a common issue where voices drift back to their original accents over time. This enhancement is particularly beneficial for dubbing applications.
Pricing for the standard API is set at $80 per million characters for v4 and $40 for Turbo, with temporary price cuts reducing these rates to $22 and $11, respectively, until October 12, 2026. In comparison, Sonic 3.6 is priced at $49 per million characters, while Gemini 3.8 Flash TTS is available for $16.49 per million characters.
The introduction of v4 follows the release of Eleven v3 just over a year ago and aims to provide voice agent developers with a solution that balances speed and expressiveness, a challenge they previously faced. ElevenLabs also emphasizes that it stores customer data in the US by default.
By Turing Wire Newsdesk · Sep 29, 2026 · How we work →
Summarised from The Decoder's original report by the Turing Wire Newsdesk. Read the original for the full story.
Source: The Decoder
