MiniMax H3
Native 2K with stereo sound already in it, from a brief that holds nine stills, three clips, and three audio tracks at once. You describe the scene — invideo's agent builds the brief.
What's new in MiniMax H3
Native 2K, not an upscale
H3 generates at 2K directly — 1440 pixels on the short edge from 16:9 through 9:16, and around 3.7 megapixels on a 21:9 frame. Nothing is enlarged after the fact, so fine detail and on-screen text hold up at full size.
Stereo sound, written with the picture
Score, dialogue, foley, and ambience are modelled together with the frames rather than generated as separate stems and mixed afterwards. It comes back in stereo, in sync, sitting in the same space as the shot.
Nine stills, three clips, three audio tracks
One brief takes up to twelve reference files across all three kinds. A character can come from a photo, the camera move from footage you already have, and the voice from a recording — in the same generation.
More than one shot per generation
H3 cuts natively, so a generation can move between angles instead of holding one. A short sequence arrives as a sequence, rather than as pieces you match up later.
Voices you bring with you
H3 transfers a voice from a reference recording, and replaces dialogue in a clip that's already generated. A character keeps the same voice across a whole project without being re-cast every time.
How MiniMax H3 compares to Hailuo 2.3
Capability | MiniMax H3 | Hailuo 2.3 |
|---|---|---|
Resolution | Native 2K, ~3.7MP at 21:9 | 1080p at 6 seconds, or 768p |
Clip length | 5–15 seconds | 6 or 10 seconds |
Frame rate | 24fps | 24fps |
Audio | Stereo, same pass as the picture | Not available |
Reference inputs | 9 images, 3 clips, 3 audio tracks — 12 files | A first frame |
Cuts within one generation | Yes | Not available |
Localized editing | Yes | Not available |
Open weights | Announced, not yet released | Not released |
How MiniMax H3 will work with invideo agent?
MiniMax H3 will run inside the same agentic workflow every model does on invideo: you direct in plain language, and the agent does the building. Here is what that will look like.
H3 will read 7,000 characters. You won't have to write them.
H3 takes very long prompts and uses the detail — blocking, lens, performance, sound. Writing that for every shot is a job in itself, so you describe the scene the way you'd pitch it, and the agent expands it into the brief the model wants.
Twelve reference slots, filled for you
A twelve-file brief is a lot to assemble by hand, over and over. The agent pulls your approved characters, locations, footage, and voices out of the project and attaches the right ones to each shot.
You decide how much runs without you
Generate and show me. Or check with me before the render. Or walk me through it. Set it once, change it whenever the project changes.
What MiniMax H3 is built to do best
Dialogue and voice work
Stereo speech, a voice carried over from a reference recording, and the option to replace a line in a clip that's already made. Dialogue stops being the part you add afterwards.
Multi-shot sequences
A generation that cuts between angles gets you a small scene rather than a single shot. Useful when an idea needs three beats and you'd rather not generate and match them separately.
On-screen text and brand work
H3 renders type and brand marks accurately enough to survive 2K, which puts lower-thirds, packaging, and UI animation inside the generation instead of composited on top.
Reusing motion you already have
Give H3 a clip and it can carry that motion onto a new subject. A camera move you liked, or a performance you already shot, becomes direction you can reuse rather than footage you're stuck with.
Helping creatives stay creative
Multiplayer mode
Collaborate in real time with live cursors to show what everyone's working on.
Storyboarding
Turn any script or idea into a shot-by-shot plan, then tweak as needed before generating.
Script writing
Write your script inside invideo, and ask an AI co-writer for help if you'd like.
Timeline editor
Picture Premiere Pro with full AI.
Build your own agents
Create custom agents to fill specific roles like cinematographer, music designer, and more.
From solo creatives to creative enterprises
World-class investors stand behind invideo.
Backed by the firms behind Stripe, Spotify, Flipkart, and ByteDance.
Pricing
Access to 200+ image, video, audio, music models including Seedance 2.5, Veo 3.1, Kling 3.0, Nano banana pro & Elevenlabs music.
Access to top stock providers like iStock, Storyblocks & more.
Model & agent prices are subject to change.
On-demand credit top-ups available.
MiniMax H3 FAQs
What is MiniMax H3?
MiniMax H3 is MiniMax's omni-modal generation model. It generates 5 to 15 seconds of native 2K video at 24fps with stereo audio produced alongside the frames, and it takes text, images, video, and audio as both input and output. One brief can hold up to nine reference images, three video clips, and three audio tracks. On invideo, you describe the scene and the agent operates the model.
Is MiniMax H3 the same as Hailuo 3.0?
Yes. H3 is the third generation of MiniMax's video line, and it goes by both names. Hailuo is the product name the earlier versions shipped under — Hailuo 02, then Hailuo 2.3. Same model, two names.
What's the difference between MiniMax H3 and Hailuo 2.3?
Resolution moves from 1080p to native 2K, and clips run 5 to 15 seconds rather than 6 to 10. Hailuo 2.3 generated no audio at all; H3 generates stereo sound with the picture. H3 also takes twelve reference files across images, video, and audio where 2.3 took a single first frame, cuts between shots inside one generation, and edits specific regions of a clip. On invideo, the agent uses whichever one suits the shot.
Does H3 really output 2K natively?
Yes. 2K is the default rather than an upscale pass — 1440 pixels on the short edge from 16:9 through 9:16, and roughly 3.7 megapixels at 21:9. A 768p tier is there for shots that don't need the resolution.
What can I feed H3 as a reference?
Up to nine images, three video clips, and three audio tracks, capped at twelve files per brief. Clips and audio run 2 to 15 seconds each. You can also set a first frame and an optional last frame, and prompts go up to 7,000 characters. On invideo, the agent assembles all of it out of your project.
Will a character hold across a whole clip?
Yes. Your approved characters, locations, and voices live in your project on invideo, and the agent attaches them to every generation. Because H3 reads stills, footage, and recordings as reference in the same brief, a face and the voice that belongs to it stay matched from shot to shot.
Can I fix one part of a clip without regenerating it?
Yes. H3 supports localized editing, so a specific region or element gets redrawn while the rest of the take stays where it was. Point the agent at what's wrong and it re-runs only that part.
How does H3 handle audio and voices?
Score, dialogue, foley, and ambience are modelled jointly with the picture and come back in stereo, already in sync. H3 can also carry a voice over from a reference recording and replace dialogue in a clip that's been generated already. There's no separate audio pass to line up.
Do I need to know how to prompt H3?
No. H3 will read a 7,000-character prompt, and the agent is the one writing it. You describe the scene in ordinary language; it handles the structure, the reference files, and the run.
When is MiniMax H3 coming to invideo?
Soon. It'll join the models the agent can reach, and you'll direct it the way you direct the rest — plain language, with your cast, locations, and voices already attached.
Does the agent take shot-by-shot control away?
No. You approve or reject every shot. The agent operates — assembles briefs, runs generations, edits regions, chains shots into a sequence — and the judgement stays yours. More automation is a setting, not the default.

