ElevenLabs Models Explained: v3, Multilingual v2, Flash, Music & More (Aug 2026)
Last updated August 7, 2026

ElevenLabs models are a family, not one voice generator: Eleven v3 (70+ languages, audio tags, expressive flagship), Multilingual v2 (long-form workhorse), Flash v2.5 (~75 ms, 40,000-char bulk model — Turbo is deprecated), plus Text to Dialogue, Voice Design, Voice Changer, SFX v2, and Eleven Music v2. All of it runs inside the invideo agent's audio and music tabs.
Updated August 2026
ElevenLabs models are a family of AI audio models — not one voice generator. As of August 2026 the lineup spans three current text-to-speech engines (Eleven v3, Multilingual v2, Flash v2.5), multi-speaker Text to Dialogue, Voice Design, a Voice Changer for speech-to-speech, Sound Effects v2, and the Eleven Music v2 song model, plus utilities like Voice Isolator. The right pick depends on the job: v3 for expressive performance, Multilingual v2 for polished long-form narration, Flash v2.5 for bulk and real-time work. This guide covers every current model, the deprecated ones stale articles still recommend, and how to direct them.
Which ElevenLabs model should you use?
The three current TTS models split cleanly by job, per the official models documentation:
| Model | Languages | Char limit / request | Latency | Best for |
|---|---|---|---|---|
| Eleven v3 | 70+ | 5,000 (~5 min audio) | Standard — not real-time | Expressive flagship: audio tags, dramatic delivery, dialogue |
| Multilingual v2 | 29 | 10,000 (~10 min) | Standard | Long-form narration workhorse: audiobooks, voiceover |
| Flash v2.5 | 32 | 40,000 (~40 min) | ~75 ms | Bulk generation, agents, real-time — at 50% lower price per character |
| Flash v2 | English only | 30,000 | ~75 ms | English-only real-time |
| Turbo v2.5 / v2 | 32 / EN | — | — | Deprecated — use Flash instead |
Two things here contradict most advice you'll read. First, Turbo is dead: many articles still recommend eleven_turbo_v2_5, but ElevenLabs now lists both Turbo models as deprecated — "outclassed by Flash," with Flash the named replacement. Second, the tiers are inverted for long-form: the cheap, fast model — Flash v2.5 — carries the biggest request limit at 40,000 characters, eight times v3's 5,000. The flagship is for performance quality, not bulk. For scripts beyond one request, request stitching (previous_text/next_text) keeps prosody continuous across chunks.
Eleven v3 reached general availability in February 2026 after a mid-2025 alpha. ElevenLabs' own GA benchmarks: users preferred the GA version in 72% of comparisons against the alpha, and text-normalization errors (currencies, chemical formulas, sports scores) fell from 15.3% to 4.9%. Those are first-party numbers — no independent quality score has been published for v3.
What is Text to Dialogue?
Text to Dialogue generates multi-speaker conversations in a single request — each turn carries its own text and voice ID, with overlapping speech and interruptions handled by the model. It runs on Eleven v3 only.
The real constraint isn't speakers — there is no speaker-count limit — it's text volume: the docs advise keeping each request under 2,000 characters, chunking and concatenating longer scenes. ElevenLabs is also unusually candid here: the feature is officially "not intended for use in real-time applications like conversational agents," and the docs concede "several generations might be required to achieve desired results." Budget retries into any dialogue workflow.
How does Voice Design work?
Voice Design creates a brand-new synthetic voice from a text description — no recording, no clone. Two models power it: eleven_ttv_v3 (70+ languages, supports audio tags) and eleven_multilingual_ttv_v2 (29 languages).
Every generation returns three preview candidates but charges you once, for the preview text's characters — effectively three-for-one sampling. The official prompt template is worth following verbatim: native language, gender and age range, audio quality level, a 2–5 word persona, 2–3 emotion adjectives, then a sentence on timbre and pacing. A guidance scale parameter trades prompt adherence against audio quality; ElevenLabs' worked examples sit at 20–40%. One documented trap: effects words like "reverb" or "phone" degrade output — describe the person, not the processing.
What does the Voice Changer do?
Voice Changer (speech-to-speech) re-voices an existing recording while preserving the performance — the docs list "whispers, laughs, cries, accents, and subtle emotional cues" as carried over from the source. That makes it an ADR-style tool: fix one flubbed word in your own read, or keep a character's voice consistent across takes and languages.
Limits: segments up to 5 minutes, an optional background-noise-removal flag, and 1,000 credits per minute of processed audio. The current models are Multilingual STS v2 (recommended even for English) and English STS v2 — there is no v3 speech-to-speech model yet.
Sound Effects v2: text to SFX
The SFX model (eleven_text_to_sound_v2) turns descriptions into Foley, cinematic design elements, game audio, and even musical one-shots ("90s hip-hop drum loop, 90 BPM"). Generations run 0.1 to 30 seconds (40 credits/second when you pin a duration), with a looping mode that covers anything longer. A Prompt Influence slider trades literal interpretation against creative variation, and non-looping effects export as 48 kHz WAV. Prompt in audio-post vocabulary — "impact," "whoosh," "braam," "glitch," "drone" — which the docs themselves teach.
Eleven Music v2: full songs, section by section
Eleven Music v2 (default since May 2026) generates vocal or instrumental tracks from 3 seconds to 5 minutes, with multilingual lyrics. What sets v2 apart is editorial control: section-by-section composition plans, inpainting that regenerates one section without touching the rest, mid-track genre transitions, and an audio reference that steers sound and tempo from a ~30-second uploaded track. Music Finetunes goes further — train a custom music model on your own non-copyrighted audio in roughly 5–10 minutes.
The licensing position matters as much as the model: ElevenLabs states the model is "trained only on licensed data and cleared for commercial use... no sync fees, no clearance delays," backed by licensing collaborations including Believe. Commercial music use starts at the Starter plan. Output is MP3 (44.1 kHz, up to 192 kbps) or WAV.
Voice Isolator, Voice Remixing — and what ElevenLabs doesn't do
Voice Isolator extracts studio-quality speech from noisy audio or video — up to 500 MB or 1 hour per file, with wide format support including MKV and WEBM. One misconception to retire: ElevenLabs has no music stem separation. The docs state the isolator "is not optimized for isolating individual music stems," and no vocals/drums/bass splitter exists anywhere in the product, whatever the "isolator" branding suggests.
Voice Remixing transforms an existing voice — gender, accent, style, pacing — while keeping it recognizable, with unlimited iterative remixing. It works on your clones, Voice Design voices, and Voice Library voices.
Instant vs Professional Voice Cloning
Instant Voice Cloning (IVC) needs just 1–2 minutes of clean audio — more than 3 minutes "yields little improvement," per the docs — and is available from the Starter plan ($6/month). You must confirm you have the right and consent to clone the voice.
Professional Voice Cloning (PVC) is the high-fidelity path: 30 minutes minimum, 2–3 hours recommended, with mandatory voice verification — you record verification lines matched against your samples before training begins. Training takes 2–6 hours, supports 38 languages, and requires the Creator plan or above. One caveat: as of August 2026 the docs flag v3 professional clones as less optimized than on Multilingual v2 and Flash v2.5, so run PVC narration on v2-family models until that changes.
How do you prompt Eleven v3 with audio tags?
v3's defining feature is inline bracketed audio tags that direct the performance:
- Emotion:
[happy],[sad],[sarcastic],[mischievously] - Delivery:
[whispers],[sighs] - Non-verbal:
[laughs],[laughs harder],[clears throat],[gulps] - Environment:
[applause],[gunshot],[explosion] - Experimental:
[strong French accent],[sings]
Example: [whispers] Someone's coming... [gulps] hide. [laughs] Just kidding — [sarcastic] you should have seen your face.
Three rules make tags land. First, set the stability mode: Creative ("emotional and expressive, but prone to hallucinations" per the docs), Natural (balanced), or Robust (stable but less responsive to tags) — use Creative or Natural when directing. Second, punctuation is a control surface: ellipses add pauses and weight, CAPS add emphasis. Third, voice selection dominates — a tag only works if the voice's training samples contain that behavior.
Pronunciation control differs by generation: v3 takes inline IPA in forward slashes (a documented 80–90% consistency rate) and does not support SSML break tags; Multilingual v2 and Flash use SSML <phoneme> tags and pronunciation dictionaries instead.
Strengths and weaknesses, honestly
Strengths. ElevenLabs is the commercial leader in AI audio — it crossed $500M ARR in May 2026 and raised a Series D at an $11B valuation in February 2026. Its differentiation is directability: audio tags, 70+ languages, multi-speaker dialogue, and an ecosystem no single-model rival matches. Zapier's June 2026 review called it the best "all-in-one voice and sound creation platform," and MasterClass's CPO/CTO Mandar Bapaye put the gap bluntly: "night and day... ElevenLabs delivered way above everyone else."
Weaknesses. Blind-preference voting is more competitive than the market position suggests: on Artificial Analysis' crowd-voted speech arena (August 2026 snapshot), Eleven v3 sits around 10th while being the joint most expensive model listed at $100 per million characters — you pay a premium for the toolchain and direction, not for an uncontested quality lead. The docs also self-report real limits: v3 isn't real-time, dialogue may need several generations, output is nondeterministic even with seeds, and Flash leaves number normalization off by default outside Enterprise. None of this dents its standing as the most complete audio stack available — it just means picking it on workflow, not leaderboard rank.
What can you make with ElevenLabs models?
- Voiceover and narration — Multilingual v2 with request stitching for long scripts: the natural engine for an AI voice generator workflow or a polished AI voiceover over edited footage.
- Your own voice at scale — IVC or PVC via AI voice cloning, then any script in your voice, in up to 38 languages.
- Localization — the family's 70+ v3 languages pair directly with AI dubbing workflows.
- Music-led video — Eleven Music v2's cleared-for-commercial tracks slot into a lyric video maker without sync-licensing anxiety.
ElevenLabs models FAQ
Can I use ElevenLabs audio commercially?
Yes, on any paid plan — the commercial license starts at Starter ($6/month), per the official billing docs.
Can I use the free plan for client or monetized work?
No. Free-tier output is licensed for non-commercial use only and requires attribution to ElevenLabs.
Whose voice can I clone?
Only voices you have the right and consent to clone — IVC requires an explicit confirmation, and PVC enforces it technically through mandatory voice verification against your own recorded samples.
Can ElevenLabs clone a celebrity voice?
No. The safety system blocks "celebrity and other high risk voices" as no-go voices, and PVC's verification step means you can only professionally clone a voice you can perform live.
Can ElevenLabs audio be detected as AI?
ElevenLabs ships a free public AI Speech Classifier to detect its own generated audio and supports C2PA content provenance — its stated principle is that "people should know when they're interacting with AI."
Is Eleven Music safe for ads and film?
Yes — ElevenLabs states Music is trained only on licensed data and output is cleared for commercial use with no sync fees, from the Starter plan up.
Which ElevenLabs model is best for long audiobooks?
Multilingual v2 for quality (10,000 characters per request plus request stitching), or Flash v2.5 when volume and cost dominate — its 40,000-character limit is the largest in the family.
Where can you use the ElevenLabs family?
The whole roster covered here — text to speech, Text to Dialogue, sound effects, Voice Design, Voice Changer, and Eleven Music — runs inside the invideo agent, surfaced through its audio and music tabs, so a script can go from v3 performance to SFX bed to a Music v2 track without leaving one project or juggling API keys. The ElevenLabs hub on invideo covers the family in the platform's context, and the voice generator and dubbing workflows above are the shortest paths to hearing these models over your footage.
Version history: Eleven v3 alpha mid-2025 → v3 GA February 2026 → ElevenMusic launch April 2026 → Music v2 (section plans, inpainting, price cuts) May 2026 → Dubbing v2 May 2026. Turbo v2 / v2.5 deprecated in favor of Flash. Facts current as of August 2026, per official ElevenLabs documentation.