Should I use a storyboard sketch or a reference photo as input for AI previz?
Last updated August 1, 2026
Use the reference photo if you must pick one — a locked reference frame scored 6/10 accuracy at 7.5/10 speed, while a raw sketch reached 7/10 accuracy but bled its art style into early generations. The best documented result converts the sketch into photorealistic frames first: 8.5/10 accuracy at 8/10 speed, $175–200 per minute.
Decide based on what each input actually does to the output, then use the hybrid that beats both. invideo is an agentic video creation tool with the current video and image models available, so the comparisons below all run through one place — the invideo agent routes your input to the right model per shot.
Feeding the sketch directly costs iterations and contaminates style. Adding hand-drawn storyboard sketches as a grid input scored 5/10 speed and 7/10 accuracy in documented testing, needed about six generations, and the sketch's art style leaked into the first video generations before it cleared. As invideo's creative team put it after testing: "It's only a little better than Workflow 2, for a lot more work." The viral social-media version of this workflow — storyboard at the bottom, generated video playing on top — was assessed as "partially or mostly" misleading about its real accuracy, so don't calibrate expectations on it.
Feeding a reference photo is faster but weaker on framing control. A locked reference frame — character, location, palette, and angle all defined in one image — scored 7.5/10 speed and 6/10 accuracy, but took 7–8 iterations to land one specific camera angle, because a photo carries look, not composition. Lock the frame before generating: even adding just a locked color-palette reference lifted accuracy from roughly 3.5/10 to 6/10 versus text-only input. Both approaches sit in the same $150–175 per minute bracket, so the sketch's accuracy gain costs you only time, not money.
The stronger answer is sketch-to-photorealistic-frames-to-video. Upload your storyboard sketch to the invideo agent, have GPT-Image-2 generate multiple candidate 3×3 image grids in your film's look, handpick the frames that match your composition, compile them into a final grid, and only then animate with Seedance 2.0. This converts your sketch's framing intent into concrete photographic references the video model can actually read — scoring 8/10 speed and 8.5/10 accuracy at $175–200 per minute, reaching a stitched final in 5–6 generations. "The frames come back looking almost like the first frames of your actual film." Two execution details matter: don't animate before you've locked and compiled the grid, and split the nine shots into three sets of three before feeding them to Seedance 2.0 — larger panels read more accurately and you keep control over edit pacing.
If you also need exact camera movement, drawing motion arrows directly on that grid pushes accuracy to 9.5/10 — but at $1,500–$2,000 per minute, roughly a 10x cost jump reserved for VFX-grade previz. For most directors and agencies, sketch-converted-to-photoreal-frames is the sweet spot: any of these tiers still lands far under traditional previz, which runs tens of thousands of dollars and takes weeks.
Watch some of these to see what works for you:
It's only a little better than Workflow 2, for a lot more work.
— invideo's creative team, on feeding raw storyboard sketches versus a locked reference image