AI Filmmaking

How do you use pre-recorded audio to drive character performance in AI video generation?

Last updated August 1, 2026

Record every line of dialogue before you generate a single shot, then feed that audio in as a generation input: upload the MP3, your script as a PDF, and a character reference image to the invideo agent, and specify each shot's dialogue inline in the prompt. The performance gets baked into the generation itself, so voice and delivery hold across every shot.

Feed the audio in before generation, not after — by default, every generated shot produces a different voice, and post-processing can't repair that because the input audio varies per generation, so even a fixed voice profile in a tool like ElevenLabs returns inconsistent output. Baking pre-recorded dialogue into the generation is the workflow that holds every time. invideo is an agentic video creation tool with the current video models available inside it, and the invideo agent routes your audio, script, and reference to a model that accepts audio as a generation input. Here is the workflow in order:

1. Record all dialogue first. Before generating any shot, record every line — your own voice, a friend's, or a hired voice actor all work. You're capturing two things at once: a consistent voice and the exact performance (pacing, emphasis, emotion) you want the character to deliver.

2. Prepare three attachments. The workflow uses exactly three file types: the dialogue as an MP3, the script as a PDF, and a character reference image. The script gives the invideo agent the scene context to match lines to shots; the image anchors who is speaking.

3. Upload all three to the invideo agent and prompt your shots. You can request multiple shots in a single prompt — one documented setup generated an establishment shot, a medium wide, and a closeup simultaneously, with scene context, shot labels, and each shot's dialogue specified inline. The invideo agent maps your recorded lines to the right shots and passes the audio through as a conditioning input.

4. Generate — the dialogue is baked in. Because the audio drives the generation rather than being overlaid in post, the character's lip movement, timing, and delivery are built around your recording from the first frame, and the same voice carries across every shot in the scene.

On model choice: this works because current-generation models accept audio as an input — Seedance 2.0's audio conditioning is the clearest example, and it's one of the more underused features in its toolset. Every roster model, including Seedance 2.0, Veo, and Kling, is available inside invideo, so the invideo agent picks the right one for the shot; you don't route audio to a model yourself.

One dependency worth knowing: audio consistency solves the voice, not the face — if your character reference image itself is inconsistent, no downstream model fixes that, so lock a clean visual reference before you start generating.

Before you start generating any single shot, take all your dialogues and record them yourself, or get a friend to record them, or get an actual voice actor to perform out the dialogues. Trust me, this is going to help a lot.

— invideo's creative team

Share

More on AI Filmmaking