Blog

Seedance 2.0 vs MiniMax-H3: Reference Systems, Arenas, and Billing Compared (Updated August 2026)

Last updated August 7, 2026

Seedance 2.0 vs MiniMax-H3: Reference Systems, Arenas, and Billing Compared (Updated August 2026)

Seedance 2.0 and MiniMax-H3 are champions of opposite arenas as of August 2026: Seedance leads with-audio image-to-video (Elo 1196), H3 sits #2 in with-audio text-to-video (1242). Seedance's @-reference system replicates a sample clip's camera and edit style and offers 4K; H3 counters with natural-language Omni-Reference, native stereo 2K, and flat $0.13/s billing. Both run side by side in invideo.

Updated August 2026

Seedance versus Hailuo usually gets framed as "China's two frontier video models," which misses the actual story: these two ship the most sophisticated reference-control systems in the market, and choosing between them is mostly choosing between those systems. As of August 2026, ByteDance's Seedance 2.0 is the image-to-video champion (#1 on the with-audio i2v board) while MiniMax-H3, released July 31, 2026, sits #2 on with-audio text-to-video, three Elo points off the overall lead. Champions of opposite arenas, priced by opposite billing models. Everything below comes from official documentation, published pricing, and dated arena data — a documented-evidence comparison, not a hands-on test — and it closes with a per-job call.

How do the @-reference and Omni-Reference systems compare?

Nobody covers this head-to-head, yet it should decide your pick. Both models accept up to 12 reference files per generation; what differs is what references can do and how you address them.

Seedance 2.0's @-reference system takes 9 images, 3 videos (≤15s each), and 3 audio files, addressed inline mid-prompt as @image1, @video1, and so on. Its signature trick is video-reference replication: hand it a sample clip and it copies that clip's camera work and editing style onto your content — the closest any model comes to "shoot it like this."

MiniMax-H3's Omni-Reference takes 9 images, 3 videos, and 3 audio files (12 files, 64MB total) but drops special syntax entirely: you reference assets in natural language — "the character in Image 2" — inside prompts that can run to 7,000 characters. H3 is omni-modal — video, image, audio, and editing in one model — so the same reference stack feeds all of it.

  Seedance 2.0 (@-references) MiniMax-H3 (Omni-Reference)
File budget 12 (9 img / 3 vid ≤15s / 3 aud) 12 (9 img / 3 vid / 3 aud), 64MB cap
Addressing Inline handles: @image1, @video1 Natural language: "the character in Image 2"
Signature capability Video-reference replication (copies camera + edit style) One reference stack across gen and editing; 7,000-char prompts
Reference cost Included in token count Video refs billed per input second; first 5 image refs free, then $0.04; audio refs free

The verdict: Seedance's system is the more cinematic instrument — style replication from a sample clip has no H3 equivalent — while H3's is the more forgiving one: no syntax to learn, and audio references cost nothing.

Where do Seedance 2.0 and MiniMax-H3 rank in the arenas?

On the Artificial Analysis with-audio boards as of August 2, 2026: text-to-video has MiniMax-H3 at #2 with an Elo of 1242 — three points behind leader Gemini Omni Flash (1245) — while Seedance 2.0 sits #3 at 1225. Image-to-video inverts it: Seedance 2.0 holds #1 at 1196, H3 #3 at 1184. H3 also leads the video-editing arena.

These rankings are a job map, not a podium. The 17-point t2v and 12-point i2v gaps are modest, drifting crowd-preference edges — but the pattern matches the design centers: Seedance grew out of image-anchored generation, H3 launched text-first.

Which has the stronger audio story?

H3 makes the louder claim: native stereo audio, generated alongside video at up to 2K. Every Hailuo model before H3 was silent — which is why Hailuo 02 and 2.3 don't appear on the with-audio arenas at all.

Seedance has generated audio natively since 1.5 Pro (December 16, 2025 — ByteDance's first audio-visual joint generation, including dialect lip-sync), and Seedance 2.0 adds lip-sync to uploaded audio: bring a voice track as an audio reference and the character mouths it. Verdict: H3 for generated-from-scratch stereo sound, Seedance for syncing to audio you already have.

How long can each generate, and at what resolution?

Length first: Seedance 2.0 generates 2–12 seconds via API (15 seconds on ByteDance's consumer apps — the ceiling differs by surface). H3 generates 4–15 seconds, so on the API, H3's 15 beats Seedance 2.0's 12. (Seedance 2.5, public since July 31, 2026, does 30 seconds — but that's the next generation, not this page's subject.)

Resolution is a philosophical split. Seedance 2.0 spans 480p–1080p and added 4K on June 23, 2026 — the same day Seedance 2.5 was announced, which is why the 4K update keeps getting misattributed to 2.5. H3 deliberately skips 4K: it targets native 2K through its H3-VAE and In-Context Regeneration, producing 2K directly rather than upscaling. Verdict: if the deliverable spec says "4K," only Seedance 2.0 checks the box; if it says "sharp, no upscaler artifacts," H3's native 2K is the more honest pipeline.

Token billing vs per-second billing: what does a clip actually cost?

The least-covered practical difference. H3 bills flat per second of output: $0.13/s at 2K. A 10-second clip is $1.30, a 15-second clip $1.95 — knowable before you hit generate. Reference billing is asymmetric in your favor: audio references free, first 5 image references free ($0.04 each after), reference video billed per input second. Failed generations aren't charged.

Seedance bills in tokens. Volcano Engine's official rates are ¥28 per million tokens for the video-input variant and ¥46/M for pure generation; BytePlus lists the 720p tier at $7/M tokens (no video input) and $4.30/M (with video input). Token consumption scales with resolution and duration, so per-clip price moves with your settings — official material puts standard-tier generation in the neighborhood of ¥1 per second, but treat that as a magnitude, not a quote. Reseller per-second Seedance rates run 2–3× above the token math — official rates only here.

Verdict: H3's flat rate wins on predictability; Seedance's token model rewards dropping resolution on drafts.

How do their guardrails and moderation differ?

Seedance 2.0's global build (shipped March 26, 2026 via Dreamina; China release February 12) carries documented guardrails: uploads of real faces are blocked, recognizable IP characters are refused, and outputs carry a visible AI label plus an invisible watermark and C2PA metadata. The China build (Jimeng/Doubao) is governed differently — don't plan a workflow around a China-app feature demo. The posture has context: Disney's cease-and-desist to ByteDance (February 13, 2026), Paramount's claims, and a March senators' letter make Seedance's clampdown a legal position, not a whim.

MiniMax's documented content posture for H3 is thinner — fewer published refusal walls, but also less certainty about where they are. The user-rights side is clearer: paid Hailuo subscribers retain output IP including commercial use, with watermark-free outputs from the Standard plan up. One promise to hold them to: MiniMax announced open weights for H3, and as of August 2026 nothing has appeared on Hugging Face. Verdict: Seedance for a documented provenance and compliance story; H3 when Seedance's face and IP blocks refuse a legitimate workflow — with the caveat that H3's limits are less documented, not absent.

Which should you use for which job?

Your job Pick Why
Image-to-video: products, ecommerce, stills brought to life Seedance 2.0 #1 with-audio i2v (Elo 1196, Aug 2026); image-anchored by design
Text-to-video with sound, no source assets MiniMax-H3 #2 with-audio t2v (1242), native stereo
Copying a reference clip's camera/edit style Seedance 2.0 Video-reference replication; no H3 equivalent
Iterating on generated footage (edit passes) MiniMax-H3 Leads the video-editing arena; omni-modal by design
4K deliverable Seedance 2.0 4K since June 23, 2026; H3 caps at native 2K deliberately
Syncing characters to an existing voice track Seedance 2.0 Lip-sync to uploaded audio references
Predictable per-clip budgeting MiniMax-H3 Flat $0.13/s vs resolution-dependent token math
Longest single API generation MiniMax-H3 15s vs Seedance 2.0's 12s API ceiling

Scan the pick column and the pattern is blunt: Seedance when you're starting from assets you have, H3 when you're starting from a prompt.

Can you run Seedance 2.0 and MiniMax-H3 on the same project?

Yes — and for this pair that's the setup that makes sense, because the honest answer above is "it depends on the shot." Both models are in the roster at invideo, the AI video platform that gives serious creatives every major model in one place, where per-shot model sequencing sends the product-stills shot to Seedance 2.0 and the from-scratch establishing shot to H3 — same project, same prompt language, character and location memory holding across both. Current status lives on the Seedance 2.0 hub and the MiniMax hub; for single-model deep dives, see the Seedance 2 guide and the Hailuo guide.

Seedance vs Hailuo, asked directly

Is Seedance better than Hailuo's H3? Per crowd-voted arenas as of August 2026: Seedance 2.0 leads image-to-video (#1 with-audio, Elo 1196); H3 leads text-to-video between the two (1242 vs 1225). Neither wins outright — they're strongest in opposite directions.

Which is cheaper, Seedance or MiniMax-H3? Not directly comparable: H3 is a flat $0.13/s at 2K; Seedance bills per million tokens (Volcano ¥28–46/M; BytePlus $4.30–7/M at 720p), so cost scales with resolution and duration. H3 is more predictable; Seedance can run cheaper at low resolutions.

Does MiniMax-H3 support 4K? No, by design — H3 targets native 2K via H3-VAE and In-Context Regeneration rather than upscaling. Seedance 2.0 is the one with 4K, added June 23, 2026.

Can Seedance really copy the style of a reference video? Yes — video-reference replication transfers a sample clip's camera work and editing style onto your generation (up to 3 video refs, ≤15s each). H3 accepts video references but documents no equivalent.

Are the MiniMax-H3 weights open source? Promised, not delivered. MiniMax announced open weights at the July 31, 2026 launch; as of August 2026 they have not appeared on Hugging Face.


Share