
Invideo today introduced Multi-Input Intelligence in Agent Two: the ability to read, watch, and understand anything a filmmaker hands it, in any format. Images, videos, scripts, PDFs, YouTube links, entire Drive folders. Agent Two takes them in the way a crew member takes a brief, and works from what it finds inside.
The problem it solves is one every filmmaker knows: the reference exists, but the tool cannot see it. The look you want is in a film you saved. The camera move you want is in footage you shot on your phone. The performance you want is in a montage of an actor's best work. Until now, all of it had to be translated into words before an AI could use any of it, and something was always lost in the translation. Multi-Input Intelligence removes the translation. Hand Agent Two the thing itself.
What makes this more than file support is what Agent Two does after it reads. Every input is checked against the project's Context, the persistent memory introduced on day one. Drop a rough cut of a short film into the project and Agent Two maps it against the shot list, reports which shots are done and which remain, and flags what a continuity supervisor would catch: a flag pole that changes between shots one and two, blood in shot seven a hotter red than the rest of the grade. It does not just watch the video. It checks the video against everything it knows about your film.
That understanding runs in both directions. Filmmakers are using it to hand over intent that never fit in a prompt box:
Patchwork and B-roll from your own footage. Upload a finished film and call for the shots you missed. Agent Two reads the lighting, the angles, and the grade of what you shot, and builds the missing shots to match, then places them where you tell it they belong.
Visual treatments pulled from references. Upload a film whose lighting you love and direct Agent Two to apply it across a sequence. Shoot the camera move you want on your phone, hand it over, and direct the agent to pull that move into the scene you are building.
Format-true product swaps. Take a winning ad, drop in new products, and direct the agent to keep everything else. Agent Two reads the ad's structure (the outfit change, the cut, the lighting shift at the midpoint), splits it accordingly, and rebuilds each half around the new product. The brief is the same one you would give a team member: here is what works, do not change the format, swap the product.
Performance direction from real acting. Upload a montage of an actor's performances and Agent Two breaks down the beats, the emotional register, and the quirks, then carries that performance style into every future generation of your character.
"The craft is still yours," said Vishal B, Resident Filmmaker, invideo. "The taste, the choice of shots, the lighting, the story: all of it comes from you. What changed is that the agent can now see what you see. I shot the camera move I wanted, uploaded it, and Agent Two understood it had to pull that move from my footage and apply it to the scene I was creating. That is how you brief a crew, and now it is how you brief the agent."
Multi-Input Intelligence builds on the foundation of Agent Intelligence: everything the agent reads becomes part of what the project knows, held in Context and checked against every generation that follows. The reference you upload today is still shaping shot four hundred.
Multi-Input Intelligence is live now for all Agent Two users at invideo.io
Part of 12 Days of invideo Agent Two
About invideo
Invideo is the AI video platform for serious creatives. Its agentic platform takes projects from script to finished film, holding characters, worlds, and creative rules across an entire production. Invideo Agent One ranks first on Physion-Arc 1.0 an independent human evaluation of seven text-to-video agents across 100 cinematic prompts, scored on 16 metrics. Invideo Agent Two, the company's frontier intelligence for creative work, powers filmmakers, studios, marketing teams, and enterprises worldwide.