Blog

Seed TTS 2.0 and Seed Audio 1.0: ByteDance's AI Voice Models Explained (2026)

Last updated August 7, 2026

Seed TTS 2.0 and Seed Audio 1.0: ByteDance's AI Voice Models Explained (2026)

ByteDance's audio line spans Seed TTS 2.0 (Oct 2025) — tone-context TTS with 200+ voices, <300ms streaming, and ~5-second Seed-ICL 2.0 cloning — and Seed Audio 1.0 (Jul 2026), which generates complete audio scenes: multi-character dialogue, SFX, ambience, and music in one ~2-minute pass across 20+ languages. Vendor-reported quality figures are labeled; both models run in the invideo agent's Audio tab.

Updated August 2026

The most interesting model in ByteDance's Seed TTS family is no longer a voice generator. Seed Audio 1.0 (July 20, 2026) generates complete audio scenes — multi-character dialogue, sound effects, ambience, and music — in a single pass: closer to an automated audio-post department than a voice generator. It sits alongside Seed TTS 2.0 (October 2025), a tone-context TTS engine that reads emotional intent from natural-language directives, and together the two make up ByteDance's audio line as of August 2026.

What models are in the ByteDance Seed audio family?

Model Released What it does Headline spec
Seed TTS 2.0 Oct 2025 Context-aware TTS with instruction-driven tone control 200+ preset voices, <300 ms streaming first packet
Seed-ICL 2.0 Oct 2025 (with 2.0) Voice cloning companion to TTS 2.0 Clones from ~5 seconds of audio
Seed Audio 1.0 Jul 20, 2026 End-to-end audio scenes: dialogue + SFX + ambience + music ~2 min per generation, 20+ languages

The original Seed-TTS research model set the direction: ByteDance's tech report described it as generating speech "virtually indistinguishable from human speech," with in-context voice cloning matching ground-truth recordings on similarity and naturalness. The 2026 family turns that research into shipping products.

What is tone-context control in Seed TTS 2.0?

Seed TTS 2.0's signature feature is that you direct the performance in plain language rather than markup. Two mechanisms stack, per a Volcengine developer-community write-up (a community article, not first-party press): global directives such as <overall emotion: angry, tone: arguing, speed: fast, pitch: high> set the scene, and inline descriptors such as [trembling urgently] steer individual lines. The same article frames the shift as "a transition from 'text reading' to 'precise emotional expression after understanding'" — and ByteDance's official Volcengine product page backs the core claim, stating the model "intelligently predicts text emotion and tone based on context."

The rest of the spec sheet, as of August 2026:

  • 200+ preset voices across styles and personas.
  • Streaming latency under 300 ms to first packet over WebSocket — fast enough for interactive use, though not the fastest in the market.
  • Seed-ICL 2.0 cloning from roughly 5 seconds of reference audio, with ByteDance claiming 97.5% average similarity, emotional renditions included — a vendor figure with no independent benchmark behind it.
  • Chinese–English bilingual synthesis with what the official page calls "natural, seamless code-switching."

One caveat: the model is billed as multilingual, but ByteDance's official English-language pages only substantiate deep coverage for Chinese and English as of August 2026 — primary distribution is Volcengine, ByteDance's China-facing cloud, and documentation depth lags outside those two languages.

What makes Seed Audio 1.0 different from a TTS model?

Everything before it in this family speaks. Seed Audio 1.0 produces. Announced July 20, 2026, it "jointly models voice, sound effects, ambience, and other audio elements within a unified framework, enabling end-to-end film-grade audio creation," per ByteDance's official Seed page. In practice: describe a scene, and one generation returns multiple characters in conversation, the room around them, the effects, and a music bed — no stitching separate renders together.

Documented capabilities, from the official Seed Audio 1.0 page (August 2026):

  • ~2 minutes per generation, with continuation support for longer pieces.
  • A structured dialogue timeline with 100 ms timing precision, so line placement and overlaps are controllable rather than left to chance.
  • Zero-shot voice creation from either a text description or a reference clip, plus cross-lingual voice transfer that preserves a speaker's rhythm, stress, and emotion in another language.
  • 20+ languages, including Chinese, English, Japanese, Korean, Spanish, Indonesian, German, French, Thai, and Vietnamese.

ByteDance self-reports an audio "availability rate" — usable-output rate — above 90% in most scenarios, and multilingual naturalness above 4.0 on a 5-point scale for most languages. Both figures are the vendor's own evaluations; no independent numbers exist yet for a model this new.

Where do Seed TTS 2.0 and Seed Audio 1.0 fall short?

Three limitations stand out. The benchmark record is thin: the 97.5% cloning similarity and >90% availability rate are self-reported, with no meaningful third-party leaderboard coverage as of August 2026. Distribution is China-first: Volcengine is the primary platform, English documentation is partial, and published pricing (new-customer bundles from ¥22.50 per 100K characters, promotional, per Volcengine) is RMB-denominated. And Seed TTS 2.0's multilingual claims beyond Chinese and English remain under-documented — for, say, German, Seed Audio 1.0's 20+ language list is the better-supported path.

What can you make with ByteDance's audio models?

The one-pass scene generation maps directly onto formats that normally require an editor and a sound library:

  • Radio drama and fiction podcasts — multi-character dialogue with ambience and score in one render, then finish the episode in an AI voice generator workflow where individual lines need replacing.
  • Localization — cross-lingual voice transfer keeps a narrator's cadence across the 20+ supported languages, a natural fit for AI dubbing pipelines.
  • Short-form video audio beds — a ready-mixed dialogue-plus-ambience track to add audio to video instead of layering stock SFX by hand.

Questions people ask about Seed TTS

Is Seed TTS the same as Doubao TTS?

Effectively yes — Seed TTS 2.0 ships in China as Doubao 语音合成 2.0 on Volcengine, ByteDance's cloud platform. Seed is the research-lab branding; Doubao is the consumer/platform branding.

How fast is Seed TTS 2.0 voice cloning?

Seed-ICL 2.0 clones a voice from about 5 seconds of reference audio, with ByteDance claiming 97.5% average similarity. The speed is documented; the similarity figure is vendor-reported and unverified independently as of August 2026.

How long can Seed Audio 1.0 generations be?

About 2 minutes per generation, with continuation support to extend a scene across multiple passes while keeping voices and ambience consistent.

Can Seed Audio 1.0 generate music and sound effects, not just speech?

Yes — that is its defining feature. It models dialogue, SFX, ambience, and music jointly in one framework and renders them as a single mixed scene, rather than as separate tracks you assemble.

What languages does Seed Audio 1.0 support?

ByteDance's official page lists 20+ languages, naming Chinese, English, Japanese, Korean, Spanish, Indonesian, German, French, Thai, and Vietnamese among them, as of its July 2026 launch.

Where can you use Seed TTS 2.0 and Seed Audio 1.0?

You don't need a Volcengine account or China-platform onboarding to run either model: both appear in the Audio tab of the invideo agent — Seed TTS 2.0 listed as multilingual TTS with tone-context control, and Seed Audio 1.0 as prompt-driven audio scenes with voices, SFX, and music together. Since invideo runs 200+ models in one place, a Seed Audio scene can land directly on a video timeline next to whatever video model generated the footage; the full lineup lives on the invideo AI models index.


Version history: Seed-TTS research report → Seed TTS 2.0 + Seed-ICL 2.0 cloning (October 2025) → Seed Audio 1.0 audio-scene model (July 20, 2026). Sourced August 2026 from ByteDance Seed and Volcengine official pages; community-sourced and vendor-reported figures labeled where used.

Share