ElevenLabs has announced v3, its next-generation voice model, launched June 21, 2026. Where previous releases optimized for raw naturalness, v3 is built around expressive control: a wider emotional range, sharper accent accuracy, and native-level pronunciation across more than 50 languages.
The v3 Model
v3 is the newest generation of the ElevenLabs text-to-speech engine, and the company positions it as a step change in how voice is directed. It is engineered for the moments speech is actually used — narration, dialog, dubbing, and product readouts — where the difference between a pleasant reading voice and a performance matters.
Emotional Range
The headline improvement is emotional range at the delivery level. v3 handles whisper-quiet passages, raised energy, and conversational nuance within a single generation rather than requiring per-segment tweaking. Creators describe it as the difference between a model that speaks text and one that performs it — a shift aimed squarely at audiobooks, character dialog, and long-form voice-over where flat, even delivery is a liability.
Accent Accuracy and Multilingual Speech
v3 also tightens accent accuracy and pushes pronunciation toward native levels across its 50-plus language footprint. For multilingual work the payoff is consistency: the same voice identity can carry a character from one language into the next without drifting into a generic international register — the core requirement for believable dubbing.
Latency and Workload
The model comes with a focus on lower latency for multilingual and multiline generation, so full-scene dubs can be produced in a single pass instead of line-by-line. That pairs with the script-to-audio workflow ElevenLabs has been widening across its products, letting teams review whole translated scenes before committing to render.
Availability
v3 is rolling out across ElevenLabs APIs, web tools, and Apps, with production-tier voice, language, and dubbing options scheduled to follow. Existing v2 projects migrate progressively, and the company says model choice will remain explicit so nothing quietly changes for existing integrations.
What This Means
The v3 launch reads as ElevenLabs’ answer to a maturing market: the frontier has moved from who sounds real to who can be directed. With emotion, accent, and multilingual control bundled into one model, ElevenLabs is betting the next competitive race is not synthetic speech that fools you — but synthetic voice you can cast.