How do you use a Blender blocking render as a camera reference for AI video?
Last updated August 1, 2026
Render a low-poly blockout from Blender — camera position, move, and stand-in geometry only — and upload it to the invideo agent as a depth-map and compositional reference. An image model like Nano Banana Pro texturizes the render into a styled frame, then Seedance 2.0 or Kling animates it, preserving your exact camera angle and character placement.
Keep the Blender scene deliberately simple: boxes and low-poly stand-ins for characters and set pieces, with the camera placed and the move keyframed exactly as you want it in the final shot. Detailed models add nothing here — the render's job is to define depth, framing, and where everything sits in the frame, and community consensus on Blender-to-AI workflows is that plain blockouts outperform detailed mannequins as references (r/comfyui).
Render the blocking pass in your film's aspect ratio — a single frame if you're locking composition, or the camera move if you're directing motion.
Upload that render to the invideo agent as a compositional and depth-map reference. invideo is an agentic video creation tool with the current video and image models available, so the same chat handles the next two steps. Instruct the invideo agent to texturize the blocking with an image model — Nano Banana Pro or GPT-Image-2 — so the low-poly geometry becomes a fully styled frame: your character replaces the stand-in, your environment replaces the boxes, while placement and angle stay pixel-accurate to the Blender camera. If you've already locked a character sheet and look-and-feel document, load them into the invideo agent first — pre-loaded context is what keeps the texturized frame consistent with the rest of your film.
Animate the texturized frame through the invideo agent: Seedance 2.0 generates up to 15 seconds per shot, Kling up to 10, and the invideo agent routes the shot and carries your context into the prompt. For a camera move longer than one generation allows, generate the move in consecutive clips and instruct the invideo agent to stitch them into one shot.
Expect to spend more credits and time per shot than with prompt-only generation — that's the trade for shot-level control — and it still lands far under traditional previz, which runs tens of thousands of dollars over weeks. The same depth-map principle also works with a photographed physical object standing in for your subject, if you want a reference without opening Blender.
Watch some of these to see what works for you:
You're going to be burning far far more credits. You're going to be spending more time, but you will gain that specific control to each shot.
— invideo's creative team