AI Filmmaking

What is the pre-record-then-generate workflow for AI video production?

Last updated August 1, 2026

The pre-record-then-generate workflow means recording every line of dialogue before generating a single video shot, then feeding that audio into the AI generation pipeline so the voice is baked into each generation rather than layered on in post. It is the workflow that solves AI audio consistency: by default, every generated shot produces a different character voice.

The pre-record-then-generate workflow is a fixed order of operations: capture all dialogue audio first, generate video second, so every shot inherits the same voice and the same performance. It exists because AI video generation has no native audio consistency — every single generation produces a different voice for the same character, and there were zero native solutions before this workflow.

Step 1 — Record all dialogue before generating anything. Take every line in your script and record it as audio: perform it yourself, have a friend record it, or hire a voice actor — all three are valid sources. What matters is that the full performance exists as fixed audio files before any generation starts, because the recording is what locks both the voice identity and the emotional delivery you want.

Step 2 — Feed the audio into the generation pipeline as an input, not an overlay. The invideo agent — an agentic video creation tool with the current generation models available inside it — accepts exactly three attachment types for this workflow: your dialogue as an MP3, your script as a PDF, and a reference image of the character. Uploading all three gives the invideo agent the voice, the words, and the face as one unified generation input.

Step 3 — Generate shots with the dialogue baked in. Prompt the invideo agent with scene context, shot labels, and the dialogue specified inline; a single prompt can produce multiple shots at once — one documented demonstration generated an establishment shot, a medium wide, and a closeup from one instruction. Because the audio drives the generation itself, the character's lip movement and performance match the recording, and the voice is identical across every shot in the project.

Why generate-then-fix fails. The common alternative — generating shots with whatever audio the model produces, then running everything through a voice-conversion tool like ElevenLabs — does not hold, because the input audio varies with every generation. Even converted into the same voice profile, the output keeps shifting, since the source performance underneath is different each time. Baking pre-recorded dialogue into the generation removes that variable entirely: the input is constant, so the output is consistent.

Before you start generating any single shot, take all your dialogues and record them yourself, or get a friend to record them, or get an actual voice actor to perform out the dialogues. Trust me, this is going to help a lot.

— invideo's creative team

Share

More on AI Filmmaking