Wan 2.7

Create cinematic 1080p clips and 4K campaign stills from a single model. Just talk to invideo’s agent and let it do the prompting for you.

Wan 2.7
Trusted by teams at
Phantom X
Wonder Studios
Google
NVIDIA
Salesforce
Visa
IBM
monday.com
Cisco
Mastercard
Meta
LinkedIn
Netflix
Ford
Hilton
Siemens
HubSpot
Phantom X
Wonder Studios
Google
NVIDIA
Salesforce
Visa
IBM
monday.com
Cisco
Mastercard
Meta
LinkedIn
Netflix
Ford
Hilton
Siemens
HubSpot

Why serious creatives choose Wan 2.7

For video

Cleaner 1080p that holds under motion

Clips run 2 to 15 seconds at full 1080p, with less flicker on skin, fabric, and moving objects than earlier versions. The detail that usually breaks the moment something moves stays put.

Seamless editing by describing the change

Tell Wan 2.7 what to fix in plain English: swap a background, warm the lighting, recolor an object, rework a character’s action or a line of dialogue. The clip gets modified instead of regenerated, so what was already working stays.

More control with first and last frame generation

Define exactly where a shot starts and where it ends, and Wan 2.7 generates coherent motion between your two anchor frames. The beat lands on the frame you chose, not near it.

Lip sync that follows your audio

Your audio track drives the performance: Wan 2.7 syncs lip movement and body motion to it during generation, not in a post pass. Change the script and the mouth follows, with the character’s vocal identity intact.

Consistent characters from up to five references

Lock a face, a body, and a voice across clips using up to five reference inputs at once: images, video, and audio together. The character who opens the project is the character who closes it.

For images

Smarter composition with Thinking Mode

Wan 2.7 Image reads the prompt, plans the composition, places the subject, sets the lighting, and checks the layout before it generates. The thinking happens first, so the output arrives coherent instead of arriving fast and wrong.

Studio quality at 4K

Stills generate at up to 4K from a text prompt, clean enough for print, web, and large format. Magazine-cover resolution without a shoot.

Exact brand colors with palette control

Enter HEX codes and the proportion each color should occupy in the frame, or pull a palette straight off a reference image. Brand colors land as specified, not as the model’s interpretation of them.

Unique faces with Thousand-Face Realism

Direct bone structure, eye shape, face shape, brow arch, nose bridge, and jawline through the prompt. Every face comes out its own face, which is how the AI same-face problem stops being yours.

Legible text across 12 languages

Wan 2.7 renders a full page of legible text in one image: signage, labels, tables, formulas, dense layouts. The words in the frame are actually words.

Style consistency from nine reference images

Feed up to nine references into one generation or edit, and Wan 2.7 Image reads them together for context-aware output. Storyboards, campaign sets, and product ranges come out looking related.


Wan 2.7 vs Wan 2.6

Feature

Wan 2.7

Wan 2.6

First & Last Frame Control

Yes

No

Prompt-Based Video Editing

Yes

No

9-grid image-to-video

Yes

No

Subject and voice cloning

Yes

No

Native audio sync

Yes

No

Text-to-image

Yes

No

Max image resolution

Upto 4K

Upto 2K

Thinking mode

Yes

No

Text rendering

3,000+ tokens in 12 languages

Basic


How Wan 2.7 works with invideo agents

On invideo, Wan 2.7 runs inside an agentic workflow. The agent routes work to it, writes the prompts, keeps your references on every generation, and edits clips instead of regenerating them, all for you to approve. Here is what that looks like in practice.

The agent picks the right model intelligently.

Wan 2.7 sits on a roster of 200+ models, and the agent routes work to it where it wins: high-volume production, shots with text that has to be legible, stills that have to hit brand colors exactly, and any project where the clips and the images should come from the same place. You never choose a model unless you want to, and if you have routing rules of your own, they hold for the whole project.

You direct, the agent writes the prompts.

Wan 2.7 plans before it generates, so what you give it to plan against decides the output. You describe the shot in plain language and the agent writes it with the structure and detail Thinking Mode builds on, instead of the thin prompt most people hand it.

Your production rides along on every generation.

Approve your character references, your voice, and your brand palette as HEX values once, and the agent saves them as the project’s memory. Every Wan call then leaves with the right faces, the right voice, and the right colors already attached, across stills and clips both, so nobody re-uploads anything at scene forty.

Continuity holds across clips.

When a scene runs past Wan’s 15-second cap, the agent sets the last frame of the finished clip as the first frame of the next one and picks the frame the new shot should land on. Wan fills the motion between them, so the cut you never wanted is a cut that never happens.

Your notes become edits, not regenerations.

Say a take is close but the lighting is cold, and the agent knows to send that to Wan’s editing path rather than spend a new generation on it. The take you approved survives the note, and the round trip costs a sentence.

You always stay in control.

You set how much the agent does on its own: let it generate everything, ask before videos, or see every prompt before it generates. That setting is yours to make, and yours to change.


Who is Wan 2.7 for?

Wan 2.7 supports controllable video and image generation across filmmaking, advertising, and design workflows. With invideo agents underneath it, each team gets the model without learning the machinery.

Filmmakers

Create anchored shots, transitions, dialogue scenes, and 1080p sequences with the cast holding across every take. Wan 2.7 helps a scene land on the exact frame it was designed for.

Microdrama producers

Build vertical episodes in native 9:16 with faces, bodies, and voices held across scenes, and dialogue edits that keep lip sync intact. Useful when a script changes late and the season still has to look and sound continuous.

Performance Ads and UGC teams

Create hooks, product moments, and creator-led variations at volume, with the same face and voice across every cut. Useful for testing cycles where the cost per iteration decides how much gets tested.

Brand and product marketers

Generate product films and 4K campaign stills with exact HEX brand colors and legible on-frame text across 12 languages. Useful for launches, packaging visuals, and localized campaigns where the color and the copy have to be right.

Professional designers and animators

Produce storyboards, campaign sets, and layouts from up to nine reference images at once, with typography, tables, and formulas rendering cleanly inside the image. Useful when the deliverable is design work, not just a picture.

Helping creatives stay creative

Multiplayer mode

Collaborate in real time with live cursors to show what everyone's working on.

RRebeccajust now
Can we push the train arrival 2 sec later? Feels rushed.
AAiko2m
Love the warmth here — keep this lighting for the reunion shot.

Storyboarding

Turn any script or idea into a shot-by-shot plan, then tweak as needed before generating.

Script writing

Write your script inside invideo, and ask an AI co-writer for help if you'd like.

Timeline editor

Picture Premiere Pro with full AI.

Build your own agents

Create custom agents to fill specific roles like cinematographer, music designer, and more.

From solo creatives to creative enterprises

Private & SecureSOC 2 & ISO AlignedGDPR Compliant

World-class investors stand behind invideo.

Backed by the firms behind Stripe, Spotify, Flipkart, and ByteDance.

Pricing

All paid plans include:

Access to 200+ image, video, audio, music models including Seedance 2.0, Veo 3.1, Kling 3.0, Nano banana pro & Elevenlabs music.

Access to top stock providers like iStock, Storyblocks & more.

Model & agent prices are subject to change.

On-demand credit top-ups available.

Wan 2.7 FAQs

What is Wan 2.7 and what can it do?

Wan 2.7 is Alibaba’s Wan model, covering video and images in one system. It generates 1080p video from text, images, or references with first and last frame control and plain-English editing, and 4K images with Thinking Mode, exact color control, and long-form text rendering. On invideo, it runs inside the agent’s roster: you direct in plain language and the agent writes the instructions.

How do I use Wan 2.7 on invideo?

Open invideo, go to Agents & Models, and select Wan 2.7 from the Video tab for clips or the Image tab for stills. Describe what you want, upload any video, image, or voice references, and the agent handles the model-facing prompt.

What is first and last frame control in Wan 2.7?

You supply the frame the shot opens on and the frame it closes on, and Wan 2.7 generates the motion connecting them. It is the difference between describing where a shot should end and deciding it.

What is 9-grid image-to-video in Wan 2.7?

You feed a 3x3 grid of your subject shot from nine angles, and Wan 2.7 reads all nine to build a full understanding of that subject before generating. The identity holds because the model has seen every side of it, not just one.

Can Wan 2.7 clone voices and faces for video?

Yes. Wan 2.7 accepts up to five reference inputs across images, video, and audio, locking a face, a body, and a voice across every clip in a project.

How does Wan 2.7 compare to Wan 2.6?

Wan 2.7 adds first and last frame control, prompt-based video editing, 9-grid image-to-video, subject and voice cloning, native audio sync, text to image, and Thinking Mode, none of which 2.6 had. Image resolution goes from 2K to 4K, and text rendering goes from basic to 3,000-plus tokens across 12 languages.

What resolution and duration does Wan 2.7 support?

Video runs 2 to 15 seconds at 720p or 1080p, in 16:9, 9:16, 1:1, 4:3, or 3:4. Images generate at up to 4K.

What is Thousand-Face Realism in Wan 2.7?

It is prompt-level control over facial structure: bone structure, eye shape, face shape, brow arch, nose bridge, jawline. You describe the face you need and get that face, instead of the generic one every AI model defaults to.

Can Wan 2.7 generate images with exact brand colors?

Yes. Enter HEX codes and the proportion each should occupy in the frame, or extract a palette from a reference image, and Wan 2.7 renders to those values rather than approximating them.

Is Wan 2.7 free on invideo?

Wan 2.7 is priced at its original API rate like all generative models on invideo. To keep using it regularly, you'll need an active plan with credits (starting at invideo's Plus tier), since credits don't roll over and unused ones reset each month.

Do I need to learn Wan 2.7’s prompting to use it on invideo?

No. The agent writes to Wan 2.7 in the structure it rewards, references and edit instructions included. You describe the shot or the still the way you would to a person, and the agent handles the model’s language.