AI Filmmaking

How do you embed in-world text — like carved signs or written notes — into AI-generated video for storytelling?

Last updated August 1, 2026

Embed in-world text by treating it as a physical prop, not an overlay: put the exact wording in quotes in your prompt, describe its carrier (carved stone, rusted sign, handwritten note), generate it as a still with a text-reliable image model like GPT-Image-2, then animate that still with image-to-video so the lettering stays legible.

Start by writing the exact text and its physical carrier into the prompt. Quote the wording verbatim and describe the object holding it — material, placement, weathering: "a carved wooden sign reading 'HIDE AFTER DUSK', weathered paint, foreground right." One documented AI short film used exactly this device — minimalist warning signage reading "They hide in the dark — HIDE AFTER DUSK" — to deliver lore and atmosphere with no dialogue at all, which also sidesteps synthetic voice performance entirely.

Generate the text object as a still image before any video. Image generations cost far less than video generations, so lock legibility at the image stage: use GPT-Image-2 for any frame containing text, since it renders lettering without distortion, while Nano Banana Pro — the stronger default for non-text imagery — can garble words and logos. Generate at high resolution (2K default, with 4K available) so the lettering survives compression and motion.

Animate the approved still with image-to-video rather than prompting text into a fresh text-to-video generation. Seedance 2.0 generates frame by frame, and dynamic camera moves like pans and whips cause it to lose detail between frames — so hold the text with a slow push-in or a static hold rather than fast movement. Budget for iteration: documented productions averaged 3 generations per usable shot, and a mangled letter counts as a failed take.

Run the whole pipeline through the invideo agent, which has all the current image and video models and routes each step to the right one — GPT-Image-2 for the text still, Seedance 2.0 for the motion pass. Save the approved text asset to project context so the same sign or note reappears identically whenever the story returns to it. The invideo agent also reads exact written details out of a loaded script: in one documented production it surfaced a precise "$57.23" figure from screenplay dialogue and carried it into the generated asset unprompted.

Use the device where it does narrative work: a carved warning establishes threat before anything appears on screen, signage anchors a location or faction, and a handwritten note advances plot in a single silent shot. Because the text is diegetic, it reads as worldbuilding rather than exposition — the audience discovers it the way a character would.

Make sure your images look right before you start making videos because image generations are cheaper than video generations. Video generations cost a lot.

— an AI filmmaker documenting an invideo agent production workflow

Share

More on AI Filmmaking