TTS Pricing Compared: ElevenLabs, MiniMax, Cartesia, Sarvam, Gemini, Seed (Aug 2026)
Last updated August 7, 2026

TTS vendors bill in three incompatible units — characters (ElevenLabs ~$50-100/M effective, MiniMax $60-100/M, Sarvam ₹30/10K, Seed promo bundles), tokens (Gemini TTS $1 in/$20 out per M, free tier), and credits (Cartesia from $5/mo) — so sticker comparisons mislead. Normalized to one hour of narration (~54K chars), computable engines span ~$1.70-$5.40. Latency claims (sub-90 to <300 ms) are vendor figures; only Cartesia has a third-party measurement (166-190 ms). Cloning splits into $1.50-3 per-voice fees vs subscription-gated with verification. Every engine runs in the invideo agent's Audio tab.
Updated August 2026
Most TTS pricing comparisons are wrong before they start, because the major voice vendors don't bill in the same unit. ElevenLabs and MiniMax price per character, Google's Gemini TTS prices per token, and Cartesia sells credit bundles — three currencies that can't be laid side by side without doing the conversion yourself. This guide does the conversion: per-character price bands, a latency ladder separating vendor claims from measured numbers, cloning economics, and a normalized one-hour-of-narration table. Every figure is official vendor pricing or a dated third-party benchmark as of August 2026 — and every engine here runs in the Audio tab of the invideo agent, the fastest way to hear them all on one script first.
What do the major TTS models cost per million characters?
For the vendors that bill in characters, the comparison is clean:
| Engine | Official rate | Per million characters | Notes |
|---|---|---|---|
| ElevenLabs Eleven v3 | Subscription credits | ~$100/M (as listed on Artificial Analysis' arena, Aug 2026) | Joint most expensive model on that board |
| ElevenLabs Flash v2.5 | 50% lower per character | ~$50/M effective | The bulk tier: 40,000-char requests, ~75 ms |
| MiniMax speech-2.8-hd | $100/M characters | $100/M | Official pay-as-you-go page |
| MiniMax speech-2.8-turbo | $60/M characters | $60/M | Same input, faster and cheaper |
| Sarvam Bulbul v3 | ₹30 per 10,000 characters (beta) | ₹3,000/M ≈ $34/M | INR-priced; v2 still listed at ₹15/10K |
| Seed TTS 2.0 | From ¥22.50 per 100,000 characters | ≈ $31/M | Promotional new-customer bundles on Volcengine — a floor, not a rate card |
Two price bands, not a continuum: the Western flagships cluster at $50–100 per million characters, while the India- and China-priced engines land in the low $30s — roughly a third to half the Western rate, with the caveats that Seed's figure is a promotional RMB bundle and Sarvam's is beta pricing that has already doubled once. Cheaper doesn't mean lesser for the intended market: Bulbul is tuned for 11 Indian languages; Seed lives on ByteDance's China-facing cloud.
Why can't you compare TTS prices directly?
Because two of the most interesting vendors don't sell characters at all.
Gemini TTS bills in tokens. Gemini 3.1 Flash TTS Preview costs $1.00 per million input text tokens and $20.00 per million output audio tokens, with batch mode at half price and a free tier — official Gemini API pricing, August 2026. The trap: output audio tokens scale with the duration of generated audio, not script length, so no fixed per-character equivalent exists — a dollars-per-million-characters number for Gemini TTS is a fabrication by construction. Run your own script on the free tier and measure.
Cartesia Sonic bills in credits. The published tiers run Free (20K credits/month), Pro ($5/month, 100K credits plus the commercial license), Startup ($49, 1.25M), Scale ($299, 8M), and Enterprise (voice-agent calls at $0.06/minute). What the tier table doesn't state is the credit-per-character schedule, so converting bundles into a per-million-character rate requires Cartesia's own credit accounting.
Even ElevenLabs' figures carry an asterisk: it sells subscription credits, so $50–100/M is an effective rate (the $100/M is Artificial Analysis' listing for Eleven v3), not a line on an invoice. The thesis of this page follows: a TTS price comparison is only as good as its unit conversions, and half the market's units don't convert.
What does one hour of narration cost on each engine?
Normalize on the job instead. At a typical 150-words-per-minute read, one hour is about 9,000 words — roughly 54,000 characters including spaces. On that assumption:
| Engine | Rate applied | ~Cost per narrated hour |
|---|---|---|
| Seed TTS 2.0 | ¥22.50/100K chars (promo bundle) | ≈ ¥12 (~$1.70) — promotional, RMB-billed |
| Sarvam Bulbul v3 | ₹30/10K chars (beta) | ₹162 (~$1.90) |
| ElevenLabs Flash v2.5 | ~$50/M effective | ~$2.70 |
| MiniMax speech-2.8-turbo | $60/M | ~$3.25 |
| MiniMax speech-2.8-hd | $100/M | ~$5.40 |
| ElevenLabs Eleven v3 | ~$100/M effective | ~$5.40 |
| Cartesia Sonic | Credit tiers from $5/mo | Not computable without the credit-per-character schedule |
| Gemini TTS | $1/M in + $20/M out (tokens) | Not computable from script length — benchmark on the free tier |
The verdict hiding in this table: the spread between the cheapest and priciest computable rows is about 3x, not the 10x sticker prices suggest — and the two rows you can't fill prove that billing-unit opacity, not price, is the real comparison problem in TTS. For a daily-publishing channel the gap between $1.90 and $5.40 an hour compounds; for a one-off brand film it's noise — pick on quality and directability instead.
Which TTS engine is fastest? The latency ladder
Latency numbers come in two species — what the vendor claims under ideal conditions, and what someone measured with a network in the loop:
| Engine | Vendor claim | Independently measured | What the claim covers |
|---|---|---|---|
| ElevenLabs Flash v2.5 | ~75 ms | — | Model latency, per ElevenLabs docs |
| Cartesia Sonic | Sub-90 ms | 166 ms median (Vapi, Jun 2026); ~190 ms median per Cartesia's own changelog | Claim is model time-to-first-audio; measurements include streaming and network |
| Maya 1 (open source) | Sub-100 ms streaming via vLLM | — | Self-hosted on a single 16 GB+ GPU — no vendor network hop, but hardware-dependent |
| MiniMax Speech | Under 250 ms end-to-end | — | Vendor announcement, Oct 2025 |
| Seed TTS 2.0 | Under 300 ms first packet | — | WebSocket streaming, vendor docs |
Cartesia is the only engine here with both a claim and a third-party measurement, and the gap is instructive: sub-90 ms marketed, 166–190 ms measured once streaming and network enter — still inside natural conversational turn-taking, but a reason to read every unmeasured row as a best case, not a budget line. If a pause breaks your product, test on your own network path.
What does voice cloning cost — and what gates it?
Cloning splits into two economic models, and the split matters more than the sticker.
Per-voice fees. MiniMax charges $1.50 per voice for Rapid Voice Cloning (~10 seconds of source audio) and $3.00 per Voice Design creation — one-off costs small enough that price never decides. The flip side: MiniMax documents no consent-verification mechanism as of August 2026, so that $1.50 is the only gate between an uploaded clip and a working clone. Cloning rights are entirely your responsibility.
Subscription-gated cloning. Cartesia includes 10-second instant cloning from the $5/month Pro plan, with professional cloning at Startup ($49) and up. ElevenLabs offers Instant Voice Cloning (1–2 minutes of audio) from the $6/month Starter plan, and Professional Voice Cloning — 30 minutes of audio minimum, 2–3 hours recommended, 2–6 hours of training — from the Creator plan, behind mandatory voice verification: you record matched verification lines before training begins, so you can only professionally clone a voice you can perform live. Verification is friction, and it is also the clearest liability shield in the cloning market.
Two engines opt out: Sarvam Bulbul offers no self-serve cloning, and Gemini TTS ships 30 presets with no cloning surface — a feature for compliance-minded teams, since a workflow that cannot produce a clone cannot produce an unauthorized one. Seed-ICL 2.0 clones from roughly 5 seconds, with a vendor-claimed, unverified 97.5% similarity.
So which engine should you price for?
Match the billing model to the workload. High-volume narration: Sarvam for Indian languages, MiniMax Turbo or ElevenLabs Flash for global scripts — the ~$2–3.25/hour band. Premium performance reads: Eleven v3 or MiniMax HD at ~$5.40/hour, where directability justifies the 2x. Real-time agents: Cartesia or Flash, budgeted at measured latency. Zero-commitment experiments: Gemini TTS's free tier and Cartesia's 20K free monthly credits — noting Cartesia's commercial license starts at Pro, not Free.
Or skip the six billing accounts: every engine priced here is selectable in the invideo agent's Audio tab, so one script can run across ElevenLabs, MiniMax, Cartesia, Seed, Sarvam, and Gemini TTS in a single project — an AI voice generator workflow carries the winner onto a video timeline — before you decide whose API pricing is worth learning in anger.
TTS pricing questions, answered
What is the cheapest TTS API in 2026?
Among engines with published per-character rates, Seed TTS 2.0's promotional Volcengine bundles (¥22.50/100K characters, ≈$31/M) and Sarvam Bulbul v3 (₹30/10K, ≈$34/M beta) sit lowest — about a third of Western flagship rates, hedges attached.
How many characters is one hour of narration?
At a typical 150-words-per-minute read, about 54,000 characters including spaces — the working assumption behind this page's per-hour table.
Is Gemini TTS cheaper than ElevenLabs?
Unanswerable as posed — Gemini bills $1/M input and $20/M output tokens, and output tokens scale with audio duration, not script length, so no fixed per-character rate exists. Benchmark your own script on the free tier.
What is the cheapest way to clone a voice?
By sticker, MiniMax's $1.50-per-voice Rapid Voice Cloning; by subscription, Cartesia's $5/month Pro (10-second clones) or ElevenLabs' $6/month Starter (1–2 minute IVC). Professional cloning costs more in audio than money: 30 minutes to 3 hours of recordings on ElevenLabs PVC, plus mandatory verification.
Are sub-100 ms TTS latency claims real?
They are vendor model-latency figures under ideal conditions. The one engine with an independent measurement — Cartesia — shows 166–190 ms in practice against a sub-90 ms claim. Read every unmeasured claim the same way.
Do free TTS tiers allow commercial use?
Generally no. ElevenLabs' free tier is licensed for non-commercial use with attribution required, and Cartesia's commercial license starts at the $5/month Pro plan. Check each vendor's current terms before shipping client work on a $0 tier.
Sources: official pricing and documentation pages from ElevenLabs, MiniMax, Cartesia, Sarvam, Google (Gemini API), and ByteDance/Volcengine; the Artificial Analysis speech arena (Aug 2026 snapshot); Vapi's latency measurements (Jun 2026). Currency conversions use approximate August 2026 exchange rates (₹≈87/$, ¥≈7.2/$). Each engine's family page is linked where it first appears above.