Blog

HappyHorse AI: Alibaba's Video Model That Won the Arena Anonymously (1.0 vs 1.1)

Last updated August 7, 2026

HappyHorse AI: Alibaba's Video Model That Won the Arena Anonymously (1.0 vs 1.1)

HappyHorse is Alibaba's closed-weight cinematic video family: 1080p/24fps clips of 3-15 seconds with audio and lip-synced dialogue generated in a single pass, in seven languages per most reporting. It debuted anonymously on the Artificial Analysis arena around April 7, 2026 and took #1 in both text-to-video and image-to-video before Alibaba was revealed as its maker. Both versions run inside the invideo agent.

Updated August 2026

HappyHorse earned its ranking before anyone knew who built it. Around April 7, 2026, an unlabeled model appeared on the Artificial Analysis video arena and, judged purely on blind head-to-head votes, climbed to #1 in both text-to-video and image-to-video; on roughly April 10, 2026, Alibaba revealed itself as the maker. The family is Alibaba's cinematic AI video line — closed-weight, generating 1080p, 24fps video with sound and lip-synced dialogue in a single pass, at lengths of 3 to 15 seconds. Two versions now exist — HappyHorse 1.0 and HappyHorse 1.1 — and both run inside the invideo agent alongside the rest of the major model lineup.

What is HappyHorse?

HappyHorse is the second video model family from Alibaba, distinct from the older, developer-oriented Wan line. Where most video models generate silent footage and leave audio to a separate pass, HappyHorse generates video and audio jointly in one pass — ambient sound, effects, and spoken dialogue with matching lip movement. Per most reporting, native lip-synced dialogue works in seven languages: English, Mandarin, Cantonese, Japanese, Korean, German, and French (source accounts differ slightly on the exact list, so treat the count as approximate).

Alibaba has leaned into the cinematic positioning: alongside the 1.1 release on June 23, 2026, it launched the HorsePower filmmaking competition with prizes up to RMB 1,000,000 for short films made with the model. Community write-ups have speculated about the architecture behind it (a unified transformer with aggressive step distillation), but Alibaba has not published confirmed details, so this guide sticks to what is documented.

What are HappyHorse's specs?

Spec HappyHorse (1.0 and 1.1)
Max resolution 1080p
Frame rate 24fps
Clip length 3–15 seconds
Aspect ratios 5 supported
Audio Joint single-pass audio + video (ambient, effects, dialogue)
Dialogue Native lip-synced speech, seven languages per most reporting
Inputs Text, image (up to 5 reference images), reference video
Editing Natural-language video editing
Weights Closed — hosted access only

Most of this table is 2026 table stakes; the audio row is not. Joint single-pass dialogue is the spec the rest of the page keeps coming back to.

HappyHorse 1.0 vs 1.1: which version should you pick?

Both versions appear in invideo's model picker as HappyHorse 1.0 and HappyHorse 1.1. The honest summary, as of August 2026: 1.1 is the newer release and the stronger audio-video model, but 1.0 still scores higher on the silent-video leaderboard — so "newer" does not automatically mean "better" for every job.

  HappyHorse 1.0 HappyHorse 1.1
Released April 2026 (anonymous arena debut ~April 7) June 23, 2026
Arena standing (Aug 2026) ~1284 Elo, #3 on text-to-video no-audio ~1151 Elo, #5 on text-to-video with-audio
Best for Silent or music-bed footage where raw visual quality wins Dialogue scenes and anything that needs synced sound
Resolution / length 1080p, 24fps, 3–15s 1080p, 24fps, 3–15s

Practical rule: if the shot needs a character to speak, use 1.1. If it's a silent cinematic shot you'll score separately, 1.0's no-audio ranking (1284 vs 1264 for 1.1 on the same board, per Artificial Analysis, August 2026) suggests it still has the edge.

How did HappyHorse debut anonymously and hit #1?

The Artificial Analysis arena shows voters two unlabeled clips generated from the same prompt and asks which is better; Elo rankings emerge from thousands of these blind votes. Around April 7, 2026, a mystery entrant began winning — and kept winning until it held #1 in both text-to-video and image-to-video. Only after the rankings were established did Alibaba confirm, around April 10, that the model was theirs. That sequencing matters: the top placement was earned before any brand halo could influence votes. VentureBeat's coverage also noted the timing context — HappyHorse rose up the charts in the same window that OpenAI discontinued the Sora consumer app (March 24, 2026, with the Sora API scheduled to sunset September 24, 2026), reshuffling the top of the leaderboard.

What makes HappyHorse different from other video models?

The moat is speech. Most video models are silent; the few with audio typically bolt it on as a second stage, which is where lip-sync drift comes from. HappyHorse generates the audio track and the video frames together in one pass, so a character's mouth is animated against the dialogue as it's generated — in, per most reporting, seven languages spanning English, Mandarin, Cantonese, Japanese, Korean, German, and French. For anyone producing dialogue-driven content in more than one market, that collapses a generate–dub–resync pipeline into a single prompt.

What can HappyHorse actually do? The four endpoints

  • Text-to-video — prompt to a finished 1080p/24fps clip with sound.
  • Image-to-video — animate from stills, with up to 5 reference images to lock character or product identity across a shot.
  • Reference-to-video — supply reference material and generate new footage that carries its subject or style.
  • Natural-language video editing — describe a change to an existing generation in plain words rather than re-rolling from scratch.

For character-consistent work, the five-reference-image ceiling is the number that matters: it is the mechanism for keeping a face or product stable across separate generations.

Where does HappyHorse rank in August 2026?

As of August 2026 on the Artificial Analysis video arena: HappyHorse 1.0 sits at roughly 1284 Elo, #3 on the text-to-video no-audio board; HappyHorse 1.1 sits at roughly 1151 Elo, #5 on the with-audio board. The April #1 placements have since been passed as newer models entered, which is normal arena churn — top-five standing in both modalities four months after launch is still an elite result. And to repeat the honest wrinkle: 1.0 outranks 1.1 on the no-audio board (1284 vs ~1264), so the version-upgrade story is not clean; 1.1's gains are concentrated on the audio side.

Documented weaknesses

  • 15-second ceiling. Longer scenes must be built from multiple clips and sequenced.
  • A/V sync degrades in complex scenes. Lip-sync is strongest on clear, front-facing dialogue; crowded or fast-cut compositions can drift, and physics artifacts appear in longer generations.
  • Text rendering. On-screen signage and titles frequently come out garbled — a near-universal weakness in current video models, and HappyHorse is no exception. Add real text in an editor afterward.

HappyHorse vs Wan: Alibaba's two video families

Alibaba now ships two distinct video lines, and they solve different problems. Wan is the developer-heritage family — it began as an open-source project (open weights end at Wan 2.2; later versions are API-only) and its current flagship, Wan 2.7, emphasizes controllability: multi-shot planning, audio-as-input conditioning, and long structured prompts. HappyHorse is the cinematic consumer line — closed from day one, tuned for out-of-the-box visual quality and native spoken dialogue rather than pipeline control. If your job is a dialogue scene or a polished single shot, start with HappyHorse; if it's a multi-shot sequence needing tight character control, Wan 2.7 is the sibling to reach for — see our Wan 2.7 complete guide for that side of the family.

How much does HappyHorse cost?

Official pricing from Alibaba Cloud's Model Studio (Bailian), as of August 2026: ¥0.9 per second at 720p (~$0.125/s) and ¥1.6 per second at 1080p (~$0.22/s), with failed generations not billed. That puts a 10-second 1080p clip with sound at roughly $2.20 — around a third of Sora 2's per-second rate at comparable settings. Third-party resellers advertise wildly different numbers; ignore them and price against Bailian.

What is HappyHorse best at? Use cases

  • Multilingual dialogue scenes. A character delivering lines in Mandarin, German, or Japanese with native lip-sync, straight from the prompt — the core scenario for an AI video generator workflow that used to require separate lip sync passes.
  • Ad localization in one pass. Generate the same spot with dialogue in each target language rather than dubbing one master — a direct shortcut for ad-making teams shipping to multiple markets, and a complement to traditional AI dubbing for footage that already exists.
  • Cinematic short-form. Sub-15-second hero shots with synced ambient sound for social, where the arena scores say HappyHorse's raw visual quality is top-three globally.

Prompting basics

HappyHorse responds well to shot-language prompts that specify dialogue explicitly. Two patterns that map to its strengths:

  • "Medium close-up, a woman in a rain-soaked Tokyo alley looks into the lens and says in Japanese: 「もう戻れない」— neon reflections, handheld, shallow depth of field, night." Quote the dialogue and name the language; the model handles the lip-sync.
  • "Product on a marble counter, slow 180-degree orbit, soft morning light, gentle ambient kitchen sounds, no dialogue." When you don't want speech, say so — otherwise the audio pass may improvise.

Keep each prompt to one shot and one action; chain clips for anything longer than 15 seconds.

Frequently asked questions

Is HappyHorse open source?

No. HappyHorse is closed-weight and hosted-only — there are no downloadable weights and no local or self-hosted deployment, unlike early versions of its sibling Wan. Access is via Alibaba's hosted endpoints or platforms that integrate them, including invideo.

Who makes HappyHorse?

Alibaba. Reporting differs on which internal unit leads it (accounts variously credit Alibaba's ATH unit and a Taotian-affiliated lab), but the company-level attribution — confirmed by Alibaba around April 10, 2026, after the anonymous arena run — is not in question.

Is happyhorseai.com or happy-horse.net the official site?

No — and this deserves a real warning. Numerous sites presenting themselves as "official HappyHorse" are third-party wrappers reselling API access, often with made-up pricing. The only first-party access points are happyhorses.io, Alibaba Cloud Model Studio (Bailian), and the Qwen app.

How long can a HappyHorse video be?

3 to 15 seconds per generation at up to 1080p/24fps. Longer pieces are assembled from multiple clips.

Does HappyHorse generate audio and dialogue?

Yes — jointly with the video in a single pass, including lip-synced spoken dialogue in seven languages per most reporting (English, Mandarin, Cantonese, Japanese, Korean, German, French). This is its defining feature.

Should I use HappyHorse 1.0 or 1.1?

Use 1.1 for anything with dialogue or sound; consider 1.0 for silent footage, where it still ranks higher on the no-audio arena board as of August 2026.

Can I use HappyHorse without an Alibaba Cloud account?

Yes. Platforms that host the model handle the Alibaba side for you — in invideo, both HappyHorse versions appear directly in the model picker.

Where can you use HappyHorse today?

HappyHorse 1.0 and 1.1 both run inside the invideo agent, which puts them next to 200+ other models — so you can generate a dialogue shot with HappyHorse, cut it against footage from other models, and finish the edit in one place. Start from the HappyHorse model hub, or go straight to a workflow like the AI video generator and pick HappyHorse from the model list.


Version history: HappyHorse 1.0 — anonymous Artificial Analysis arena debut ~April 7, 2026; revealed as Alibaba ~April 10, 2026. HappyHorse 1.1 — released June 23, 2026, alongside the HorsePower filmmaking competition (prizes to RMB 1M).

Share