What is a depth-map reference and how is it used in AI video production?
Last updated August 1, 2026
A depth-map reference is an image that encodes spatial structure — where every object sits and how far it is from camera — so an AI model inherits composition, character placement, and camera angle from geometry instead of a text prompt. In AI video production, you build one from a low-poly Blender blocking render or even a photographed physical object, then have the model texturize that geometry into finished frames.
Use a depth-map reference when you need exact control over character placement and camera angle — the two things a text prompt pins down least reliably. Conventionally, a depth map is a grayscale image where brightness encodes distance from camera; in practice, any image with clear spatial geometry can serve as one, because the AI reads the arrangement of shapes as a 3D scaffold and keeps your framing intact while replacing the contents. invideo is an agentic video creation tool with all the current image and video models available, and the invideo agent accepts depth-map references directly as generation inputs.
There are two documented ways to create one:
A low-poly Blender blocking render. Block your scene in Blender with rough, untextured geometry, render it from your chosen camera angle, and upload the render to the invideo agent with an instruction to texturize it using an image model — Nano Banana or GPT-Image-2 both handle this step. The blocking render carries the composition and depth information; the image model dresses it in your film's look. This gives you precise control over where characters stand and exactly where the camera sits — decisions locked in 3D before any generation runs. One documented previs workspace for a monster-feature sequence ran this approach across 11 sub-agents.
A photographed physical object. You don't need 3D software at all: frame any object with real depth — a toy on a desk works — photograph it with your phone, and upload the photo as your reference. The invideo agent reads the photo as a depth map, replacing the toy with your character and the desk with your film's environment while preserving the spatial arrangement and camera perspective of your photo.
Once the texturized frame matches your intent, the invideo agent animates it through a video model — Seedance 2.0 generates up to 15 seconds per shot, Kling up to 10 — and because the geometry was fixed before generation, the composition holds instead of drifting between takes. Expect the trade-off the documented workflows describe: depth-referenced generation burns more credits and time than prompt-only generation, but you gain shot-specific control. The cost still lands far below traditional methods — documented AI previs runs came in around $150–$175 per minute of output, against previz budgets that traditionally reach tens of thousands of dollars and take weeks, with complex sequences historically costing $50,000–$100,000 for previs alone.
One adjacent technique worth knowing: a depth-map reference controls composition and framing, not motion — for directing camera movement, drawing directional arrows on a storyboard grid serves the equivalent role and is covered as its own workflow.
Watch some of these to see what works for you:
I could possibly take any of my kids' toys and frame that on my table, take an image with my phone, upload that onto Agent One, and watch it use the toy as a depth map reference — replacing the toy with your character and the desk with the environment in your film.
— Hridaye, invideo's creative director