Ranked #2 Music & Voice — Sound from Scratch
ElevenLabs

ElevenLabs v4

ElevenLabs v3 could act a line if you told it how. Eleven v4 reads the whole scene — who is speaking, what just happened, how tense the conversation has become — and performs it. Listeners in Artificial Analysis's blind voice arena now rank it first, and a faster Turbo version brings the same voices to live conversation.

Updated September 29, 2026 VoiceTTS90+ Languages
9.1out of 10
Official Website
Best for

ElevenLabs v3 could act a line if you told it how. Eleven v4 reads the whole scene — who is speaking, what just happened, how tense the conversation has become — and performs it. Listeners in Artificial Analysis's blind voice arena now rank it first, and a faster Turbo version brings the same voices to live conversation.

Why It Wins

#1 on the Artificial Analysis Provider Voice arena in September 2026 (1319 Elo, ahead of Cartesia Sonic 3.6 at 1276 and Gemini 3.8 Flash TTS at 1267). 90+ languages, up to 10,000 characters per request, stackable audio tags, IPA pronunciation, and instant voice clones from about 10 seconds of audio. The Turbo model answers in about 150 ms, per ElevenLabs.

Watch out

Listed at $0.08 per 1,000 characters ($80 per million), it costs more than Cartesia ($49) and far more than Gemini 3.8 Flash TTS ($16.49). A 72% launch discount runs until October 12, 2026. SSML break tags are not supported, and voices this realistic keep consent and misuse questions front and center.

01

What It Actually Is

There is a difference between reading a script and performing one.

Hand a line like “Oh, great, you’re here” to a text reader and it says the words. Hand it to an actor and they ask what happened just before. Is the character relieved? Sarcastic? Did someone just walk in late to a meeting?

Earlier versions of ElevenLabs could perform a line if you told them how. Eleven v4 tries to work out how on its own. ElevenLabs says it reads the surrounding conversation — who is speaking, how the exchange has escalated, what was just said — and delivers each line to fit.

ElevenLabs released Eleven v4 on September 28, 2026, alongside Eleven v4 Turbo, a faster version for real-time voice agents. Both are live in the ElevenLabs apps and API, including on the free tier.

What listeners say

The number that matters most for a voice model is simple: which one do people prefer when they can’t see the label?

Artificial Analysis runs a Provider Voice arena that answers exactly that. Listeners hear two anonymous clips of the same text and pick the more natural one, and the votes are turned into Elo ratings, like chess. In September 2026, Eleven v4 took first place with 1319 Elo, ahead of Cartesia Sonic 3.6 at 1276 and Google’s Gemini 3.8 Flash TTS at 1267. Artificial Analysis also ranks it first on pronunciation robustness — getting tricky words right — and second on controlled voice.

ElevenLabs adds its own blind test, in which about 75% of listeners preferred v4. That is a vendor-run figure, so the arena ranking is the one to rely on.

What’s new since v3

v3 introduced audio tags: short directions in square brackets, like [whispers] or [laughs], that change how a line is delivered. v4 builds on that idea:

  • Tags can be stacked and sequenced. You can combine directions, and the model follows them in order across a passage.
  • It listens to context. Replies in a dialogue take their tone from what came before, which matters for arguments, escalations, and the “please hold” moments of customer service.
  • Pronunciation you can spell out. Wrap a word in IPA, the phonetic alphabet, between forward slashes and v4 will say it that way. ElevenLabs says this is now much more reliable.
  • Longer requests. Up to 10,000 characters per request, roughly ten minutes of audio, and better stitching when you chain requests together.
  • Steadier voices. A new method keeps a speaker’s identity consistent across long narration and repeated takes, which used to be a weak spot on long projects.

Languages grow from about 70 to more than 90, with the biggest quality jumps in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. One voice can now speak any supported language while keeping its identity.

Voice cloning got faster too. An Instant Voice Clone now needs only about 10 seconds of audio. For the highest fidelity, Professional Voice Clones are supported as well.

Turbo: voices you can talk to

A voice agent lives or dies on timing. If the reply starts a full second after you stop talking, the conversation feels like a bad phone line.

Eleven v4 Turbo is built for that job. ElevenLabs reports a median inference time of about 100 ms and about 150 ms to first speech, against 262–814 ms for the competitors it tested. These are the company’s own measurements and exclude network delays, but they put Turbo inside the length of a natural pause between two people talking. It costs half as much as the full model.

The honest catch

It is the premium option. At list price, Eleven v4 costs $0.08 per 1,000 characters, which Artificial Analysis records as $80 per million characters. Cartesia Sonic 3.6 costs $49 and Gemini 3.8 Flash TTS $16.49 for the same amount. For bulk narration where “clear and pleasant” is enough, those are far cheaper. ElevenLabs is running a 72% launch discount until October 12, 2026, which brings v4 to $0.022 per 1,000 characters, but plan budgets on the list price.

The controls are different. Like v3, v4 does not support SSML break tags, the markup many older text-to-speech scripts use for pauses. Direction comes from punctuation, audio tags, and IPA instead. If you have an existing production pipeline, rewrite a few scripts and re-listen to your voices on v4 before switching everything over.

Great clones cut both ways. A convincing copy of someone’s voice from ten seconds of audio is useful for creators and dangerous in the wrong hands. ElevenLabs applies verification and safeguards, but consent and commercial rights stay your responsibility.

Who should use it

If you need… Reach for Why
Audiobooks, characters, dubbing where delivery matters Eleven v4 Top of the independent listening arena, with fine control
A live voice agent or assistant Eleven v4 Turbo ~150 ms to first speech, half the price
Huge volumes of plain narration on a budget Gemini 3.8 Flash TTS or Cartesia Sonic 3.6 A fraction of the price per character
Songs with vocals and instruments Suno v6 A music generator, not a voice tool

The verdict

Eleven v4 is the voice model to choose when the take has to sound directed rather than read. It leads the independent arena, speaks more than 90 languages, and gives you real tools to shape a performance. It is not the cheapest way to turn text into speech, and it does not need to be: when the voice is the product, it is the one listeners pick.

02

Strengths and honest limitations

Key Strengths

  • The voice listeners prefer: In Artificial Analysis’s Provider Voice arena, where people pick the more natural of two anonymous clips, Eleven v4 took first place in September 2026 with 1319 Elo. Cartesia Sonic 3.6 scored 1276 and Gemini 3.8 Flash TTS 1267. Artificial Analysis also ranks it first on pronunciation robustness.
  • Direction, not just dictation: Audio tags such as [whispers], [laughs], or [said angrily] can now be stacked and followed in sequence, and the model takes the surrounding conversation into account, so a reply sounds like a reply rather than an isolated line.
  • Built for long work: Each request can hold up to 10,000 characters (about ten minutes of audio), and ElevenLabs says a new method keeps a speaker’s identity steady across long narration and re-takes.
  • 90+ languages from one voice: Coverage grows from about 70 to more than 90 languages, with the biggest improvements in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. One voice can speak any supported language while keeping its identity.
  • Fast enough to talk to: Eleven v4 Turbo is built for voice agents. ElevenLabs reports about 100 ms median inference and about 150 ms to first speech, against 262–814 ms for the rivals it tested.

Honest Limitations

  • Premium price: At list price, $80 per million characters is 1.6 times Cartesia Sonic 3.6 and nearly five times Gemini 3.8 Flash TTS. The 72% launch discount ($0.022 per 1,000 characters) ends on October 12, 2026.
  • Vendor numbers for some claims: The 75% blind-test preference and the latency comparison come from ElevenLabs’s own tests. The arena ranking is the independent number to trust.
  • New controls to learn: Like v3, v4 ignores SSML break tags. Pauses and emphasis come from punctuation, audio tags, and IPA spellings, so older scripts may need rewriting — and it is worth re-listening to your existing voices on v4 before switching production work.
  • Clones raise real questions: A convincing clone from ten seconds of audio is powerful and easy to misuse. ElevenLabs applies verification and safeguards, but consent and commercial rights remain your responsibility.
03

Benchmark Snapshot

Artificial Analysis Provider Voice arena — #1, 1319 Elo

Independent blind listening arena, September 2026. Cartesia Sonic 3.6 1276, Gemini 3.8 Flash TTS 1267. Eleven v4 also ranks first on Artificial Analysis's pronunciation robustness test and second on controlled voice.

Blind preference — ~75% (vendor)

ElevenLabs reports that about 75% of listeners preferred v4 in head-to-head blind tests. The test was run by ElevenLabs.

Latency — ~150 ms to first speech (Turbo, vendor)

Median figures reported by ElevenLabs for Eleven v4 Turbo: about 100 ms inference, about 150 ms to first speech, excluding network and application time.

Price — $0.08 / 1K characters (Turbo $0.04)

List API prices; 72% off until October 12, 2026. Artificial Analysis lists $80 per million characters, versus $49 for Cartesia Sonic 3.6 and $16.49 for Gemini 3.8 Flash TTS. The free tier includes 10,000 v4 characters; paid plans start at $6 per month.

04

The Verdict

Eleven v4 is the voice model to pick when the delivery has to sound directed: audiobooks, game characters, dubbing, and voice agents that should sound like people rather than phone menus. It leads the independent listening arena and adds real control through stackable tags and IPA. It is not the cheapest way to turn text into speech — Gemini and Cartesia cost far less for bulk narration where good enough is good enough — but when the take matters, it is the one listeners prefer.

05

Frequently Asked Questions