Split-frame compositing is generating the left and right halves of a dialogue scene as separate AI video clips — one speaker per generation — then stitching them into a single frame in post to simulate two people talking. It's a workaround for a real model limitation: no current AI video model reliably generates multiple people speaking in the same frame.
Split-frame compositing works like this: you generate one clip of speaker A on the left side of the frame, a second clip of speaker B on the right, then combine the two halves in an editor so the final shot looks like a native two-person conversation. The result can pass casually, but the model never actually rendered a multi-person dialogue — each half was a single-speaker generation with no shared lighting, eyelines, or timing between the two.
The technique exists because multi-speaker generation is a documented cross-model failure. In one evaluation of 30+ outputs from Google's Omni Flash model, multiple people speaking in the same frame remained a persistent weakness — and the same held across every AI video model tested, not just Omni. Even single-speaker clips have a ceiling: lip sync stays consistent for roughly 6–7 seconds before degrading, so a composited two-speaker shot inherits two independent lip-sync clocks that drift on their own schedules.
Among practitioners, split-frame compositing is explicitly classified as a workaround, not a model capability. When you see a demo of a two-person dialogue built this way, it demonstrates editing skill, not that the model can hold two speaking characters in one frame. The distinction matters when you're evaluating models: a genuine multi-person generation would keep both faces, voices, and eyelines coherent inside a single output, and no current model does that reliably.
If you need a dialogue scene today, the practical route is conventional coverage rather than a composited frame: generate each speaker as a single-person clip within the 6–7 second lip-sync window and cut between them in the edit. Inside invideo, the invideo agent can route each of those shots to whichever model suits it — Veo, Kling, or Seedance 2.0 — so you build the conversation shot by shot instead of masking the limitation inside one faked frame.
Watch some of these to see what works for you:
please don't give me this I generated the left side of the frame and the right side of the frame and composited. That's just cheating.
— invideo's creative team, on evaluating AI video model capabilities