There is a difference between reading a script and performing one.
Hand a line like “Oh, great, you’re here” to a text reader and it says the words. Hand it to an actor and they ask what happened just before. Is the character relieved? Sarcastic? Did someone just walk in late to a meeting?
Earlier versions of ElevenLabs could perform a line if you told them how. Eleven v4 tries to work out how on its own. ElevenLabs says it reads the surrounding conversation — who is speaking, how the exchange has escalated, what was just said — and delivers each line to fit.
ElevenLabs released Eleven v4 on September 28, 2026, alongside Eleven v4 Turbo, a faster version for real-time voice agents. Both are live in the ElevenLabs apps and API, including on the free tier.
What listeners say
The number that matters most for a voice model is simple: which one do people prefer when they can’t see the label?
Artificial Analysis runs a Provider Voice arena that answers exactly that. Listeners hear two anonymous clips of the same text and pick the more natural one, and the votes are turned into Elo ratings, like chess. In September 2026, Eleven v4 took first place with 1319 Elo, ahead of Cartesia Sonic 3.6 at 1276 and Google’s Gemini 3.8 Flash TTS at 1267. Artificial Analysis also ranks it first on pronunciation robustness — getting tricky words right — and second on controlled voice.
ElevenLabs adds its own blind test, in which about 75% of listeners preferred v4. That is a vendor-run figure, so the arena ranking is the one to rely on.
What’s new since v3
v3 introduced audio tags: short directions in square brackets, like [whispers] or [laughs], that change how a line is delivered. v4 builds on that idea:
- Tags can be stacked and sequenced. You can combine directions, and the model follows them in order across a passage.
- It listens to context. Replies in a dialogue take their tone from what came before, which matters for arguments, escalations, and the “please hold” moments of customer service.
- Pronunciation you can spell out. Wrap a word in IPA, the phonetic alphabet, between forward slashes and v4 will say it that way. ElevenLabs says this is now much more reliable.
- Longer requests. Up to 10,000 characters per request, roughly ten minutes of audio, and better stitching when you chain requests together.
- Steadier voices. A new method keeps a speaker’s identity consistent across long narration and repeated takes, which used to be a weak spot on long projects.
Languages grow from about 70 to more than 90, with the biggest quality jumps in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. One voice can now speak any supported language while keeping its identity.
Voice cloning got faster too. An Instant Voice Clone now needs only about 10 seconds of audio. For the highest fidelity, Professional Voice Clones are supported as well.
Turbo: voices you can talk to
A voice agent lives or dies on timing. If the reply starts a full second after you stop talking, the conversation feels like a bad phone line.
Eleven v4 Turbo is built for that job. ElevenLabs reports a median inference time of about 100 ms and about 150 ms to first speech, against 262–814 ms for the competitors it tested. These are the company’s own measurements and exclude network delays, but they put Turbo inside the length of a natural pause between two people talking. It costs half as much as the full model.
The honest catch
It is the premium option. At list price, Eleven v4 costs $0.08 per 1,000 characters, which Artificial Analysis records as $80 per million characters. Cartesia Sonic 3.6 costs $49 and Gemini 3.8 Flash TTS $16.49 for the same amount. For bulk narration where “clear and pleasant” is enough, those are far cheaper. ElevenLabs is running a 72% launch discount until October 12, 2026, which brings v4 to $0.022 per 1,000 characters, but plan budgets on the list price.
The controls are different. Like v3, v4 does not support SSML break tags, the markup many older text-to-speech scripts use for pauses. Direction comes from punctuation, audio tags, and IPA instead. If you have an existing production pipeline, rewrite a few scripts and re-listen to your voices on v4 before switching everything over.
Great clones cut both ways. A convincing copy of someone’s voice from ten seconds of audio is useful for creators and dangerous in the wrong hands. ElevenLabs applies verification and safeguards, but consent and commercial rights stay your responsibility.
Who should use it
| If you need… | Reach for | Why |
|---|---|---|
| Audiobooks, characters, dubbing where delivery matters | Eleven v4 | Top of the independent listening arena, with fine control |
| A live voice agent or assistant | Eleven v4 Turbo | ~150 ms to first speech, half the price |
| Huge volumes of plain narration on a budget | Gemini 3.8 Flash TTS or Cartesia Sonic 3.6 | A fraction of the price per character |
| Songs with vocals and instruments | Suno v6 | A music generator, not a voice tool |
The verdict
Eleven v4 is the voice model to choose when the take has to sound directed rather than read. It leads the independent arena, speaks more than 90 languages, and gives you real tools to shape a performance. It is not the cheapest way to turn text into speech, and it does not need to be: when the voice is the product, it is the one listeners pick.