Blog

MiniMax Speech 2.8, 2.6, and Voice Design: The Full Audio Family (2026)

Last updated August 7, 2026

MiniMax Speech 2.8, 2.6, and Voice Design: The Full Audio Family (2026)

MiniMax Speech spans the current 2.8 flagship (native sound tags like (laughs), ~10-second voice cloning, cleaner pipeline, Jan 2026), the legacy 2.6 (sub-250 ms latency claim, Fluent LoRA cloning, Oct 2025), and Voice Design (text description to reusable voice, $3). Official pricing: $100/M characters HD, $60/M Turbo. No public cloning-consent mechanism is documented — users bear responsibility. All three run in invideo's Audio tab.

Updated August 2026

There is only one right answer in MiniMax's speech lineup as of August 2026, and it is Speech 2.8. The text-to-speech arm of MiniMax — the lab behind the Hailuo video models — covers three units: Speech 2.8 (the current flagship, in HD and Turbo tiers), the now-legacy Speech 2.6, and Voice Design, which turns a written description into a reusable synthetic voice. What settles the choice: 2.8 adds native sound tags like (laughs) and (sighs), clones a voice from roughly ten seconds of audio, and cuts background noise with a re-engineered pipeline. Everything below comes from MiniMax's official documentation and dated announcements.

Which MiniMax Speech model should you use?

MiniMax's platform documentation lists two current models and marks the 2.6 pair as legacy:

Model Released Official positioning Status
speech-2.8-hd Jan 23, 2026 "Ultra-realistic quality featuring sound tags" Current flagship
speech-2.8-turbo Jan 23, 2026 "Seamless speed meets natural flow" Current fast tier
speech-2.6-hd / 2.6-turbo Oct 30, 2025 Real-time voice, low latency Legacy

Speech 2.6 remains the family's real-time foundation. Its launch announcement (October 30, 2025) claimed "end-to-end latency of under 250 milliseconds," introduced native handling of URLs, emails, phone numbers, dates, and currency amounts without preprocessing, and shipped Fluent LoRA — cloning that preserves a speaker's timbre while producing fluent output even from imperfect, disfluent source audio, across 40+ languages. All of it carried forward into 2.8.

What's new in MiniMax Speech 2.8?

Announced January 23, 2026, Speech 2.8 makes three documented moves.

Native sound tags. You can write paralinguistic events directly into the script — the API reference supports (laughs), (chuckle), (coughs), (clear-throat), (sighs), (sneezes) and similar tags, and they are a 2.8-only feature. Instead of splicing a laugh in post, the model performs it in place.

Ten-second voice cloning. MiniMax's announcement claims fast cloning from around 10 seconds of audio: "Speech 2.8 precisely captures your unique texture, breathiness, and even your specific speaking pace." That's a first-party claim — no independent cloning benchmark for 2.8 has been published on MiniMax's pages as of August 2026.

A cleaner pipeline. The audio stack was re-engineered to reduce background noise and synthetic distortion, with fixes for cross-lingual "accent bleed" starting with Mandarin–Japanese pairs.

What are MiniMax Speech's specs?

Per the official T2A API reference, as of August 2026:

Spec Detail
Languages ~40 via language_boost (Chinese, Cantonese, English, Arabic, Hindi, Japanese, Korean, Vietnamese, Tamil, and more), plus auto-detect
Sample rates 8 kHz – 44.1 kHz
Formats MP3, PCM, FLAC, WAV, G.711, Opus
Emotion parameter happy, sad, angry, fearful, disgusted, surprised, calm, fluent, whisper
Voice shaping Pitch, intensity, and timbre controls on a −100 to 100 scale
Effects spacious_echo, lofi_telephone, robotic
Voice mixing Blend multiple voices via timbre_weights

The tell is the bottom half of that table: pitch, intensity, and timbre on a −100 to 100 scale, effects presets, and voice mixing read more like a channel strip than a typical TTS parameter list — the family assumes you will shape voices, not just pick them.

What is MiniMax Voice Design?

Voice Design is a text-prompt-to-voice API: describe the voice you want in natural language, supply preview text of up to 500 characters, and the endpoint returns a reusable voice_id plus trial audio. It's the route to an original synthetic voice — one that belongs to no real person — which you can then drive through Speech 2.8 with sound tags, emotion parameters, and effects.

How much does MiniMax Speech cost?

MiniMax publishes pay-as-you-go pricing (August 2026):

Item Official price
speech-2.8-hd (and 2.6-hd) $100 per million characters
speech-2.8-turbo (and 2.6-turbo) $60 per million characters
Rapid Voice Cloning $1.50 per voice
Voice Design $3.00 per voice (+ $30/M characters for preview audio)

That puts the HD tier at the premium end of per-character TTS pricing, with Turbo as the volume option — and one-off cloning fees low enough that per-voice cost is rarely the deciding factor.

Who uses MiniMax Speech — and the consent question

MiniMax's own Speech 2.6 announcement names LiveKit, Pipecat, and Vapi as shipping MiniMax Speech inside their voice-agent stacks — treat that as vendor-claimed adoption rather than independent confirmation, but it does signal where the family is aimed: low-latency conversational audio.

On cloning consent, state the position plainly: as of August 2026, MiniMax's public documentation describes no consent-verification mechanism for voice cloning — no enforced voice-verification step, no documented celebrity-voice blocklist. Responsibility for having the right to clone a voice sits with the user under the platform terms; if you clone anyone but yourself, get written consent first, because the API will not check for you.

Two more caveats: no independent quality benchmarks for 2.8 appear on MiniMax's official pages, and cloning is API-gated with per-voice fees rather than bundled into a subscription.

What can you make with MiniMax Speech?

  • Narration with performed emotion — sound tags plus the emotion parameter make 2.8 a strong engine behind an AI voice generator workflow, where a script can carry its own laughs and sighs.
  • Voiceover at volume — the $60/M-character Turbo tier suits long scripts run through an AI voiceover pass over edited footage.
  • Localization — the ~40-language language_boost list, with Fluent LoRA keeping timbre across languages, maps directly onto AI dubbing work.

Common questions about MiniMax Speech

Is MiniMax Speech 2.8 better than Speech 2.6?

For almost every job, yes — 2.8 adds sound tags (a 2.8-only feature), ~10-second cloning, and a cleaner pipeline, and MiniMax now marks 2.6 as legacy while 2.8 inherits its capabilities.

How much audio does MiniMax voice cloning need?

Around 10 seconds, per MiniMax's Speech 2.8 announcement (January 2026). Rapid Voice Cloning is priced at $1.50 per voice on the official pay-as-you-go page.

How many languages does MiniMax Speech support?

Roughly 40 via the language_boost parameter, plus automatic language detection, per the official API reference (August 2026).

How much does MiniMax Speech cost per million characters?

Official pricing: $100/M characters for the HD models, $60/M for Turbo. Voice Design costs $3.00 per created voice.

Does MiniMax verify consent before cloning a voice?

No public verification mechanism is documented as of August 2026. Users bear responsibility for cloning rights under the platform terms — secure consent yourself before cloning any real person's voice.

Can I use MiniMax Speech commercially?

The models are sold as a standard paid API under MiniMax's platform terms; there is no separate non-commercial tier documented on the pricing page. Review the current terms for your use case before shipping client work.

Where can you use MiniMax Speech?

Inside invideo, the MiniMax audio family sits in the agent's Audio tab — the picker lists Minimax Speech 2.8, Speech 2.6, and MiniMax Voice Design — so a script written in one project can be voiced, tagged, and cut against footage without touching an API key. MiniMax is also one of the few labs with siblings across three modalities on the platform: the MiniMax hub on invideo covers the Hailuo video models and MiniMax's music generation alongside the speech family, which matters when one brand voice needs to span narration, video, and score.


Version history: Speech 2.6 (sub-250 ms latency claim, Fluent LoRA cloning) — October 30, 2025 → Speech 2.8 (native sound tags, ~10 s cloning, re-engineered pipeline) — January 23, 2026 → 2.6 pair marked legacy in the platform docs. All facts as of August 2026, from MiniMax's official documentation and announcements.

Share