AI VFX

What is split-frame compositing in AI video editing?

Last updated August 1, 2026

Split-frame compositing is generating the left and right halves of a dialogue scene as separate AI video clips — one speaker per generation — then stitching them into a single frame in post to simulate two people talking. It's a workaround for a real model limitation: no current AI video model reliably generates multiple people speaking in the same frame.

Split-frame compositing works like this: you generate one clip of speaker A on the left side of the frame, a second clip of speaker B on the right, then combine the two halves in an editor so the final shot looks like a native two-person conversation. The result can pass casually, but the model never actually rendered a multi-person dialogue — each half was a single-speaker generation with no shared lighting, eyelines, or timing between the two.

The technique exists because multi-speaker generation is a documented cross-model failure. In one evaluation of 30+ outputs from Google's Omni Flash model, multiple people speaking in the same frame remained a persistent weakness — and the same held across every AI video model tested, not just Omni. Even single-speaker clips have a ceiling: lip sync stays consistent for roughly 6–7 seconds before degrading, so a composited two-speaker shot inherits two independent lip-sync clocks that drift on their own schedules.

Among practitioners, split-frame compositing is explicitly classified as a workaround, not a model capability. When you see a demo of a two-person dialogue built this way, it demonstrates editing skill, not that the model can hold two speaking characters in one frame. The distinction matters when you're evaluating models: a genuine multi-person generation would keep both faces, voices, and eyelines coherent inside a single output, and no current model does that reliably.

If you need a dialogue scene today, the practical route is conventional coverage rather than a composited frame: generate each speaker as a single-person clip within the 6–7 second lip-sync window and cut between them in the edit. Inside invideo, the invideo agent can route each of those shots to whichever model suits it — Veo, Kling, or Seedance 2.0 — so you build the conversation shot by shot instead of masking the limitation inside one faked frame.

Watch some of these to see what works for you:

Hands-on breakdown of Veo Omni Flash — including multi-speaker lip-sync limits

please don't give me this I generated the left side of the frame and the right side of the frame and composited. That's just cheating.

— invideo's creative team, on evaluating AI video model capabilities

Share

More on AI VFX