How do you prompt AI to generate ultra-wide cinematic framing and shot composition?
Last updated August 1, 2026
Prompt ultra-wide cinematic framing by naming the ratio AND the lens in technical terms, then stacking shot type, subject placement, depth layers, camera motion, and a negative-prompt line. A working template: "extreme wide shot, 2.39:1 anamorphic, 21mm lens, low horizon, subject lower-right third, foreground silhouette into deep midground, slow dolly-in — no centered framing, no shallow DOF, no telephoto compression."
Start by writing the frame in parameters the model recognizes, not adjectives. Lead the prompt with the ratio and the lens — "2.39:1 anamorphic", "21:9 cinematic letterbox", "ultra-wide 16mm", "21mm rectilinear" — because models latch onto numeric ratio notation and focal-length tokens far more reliably than words like "cinematic" or "epic". Pair that with an explicit shot type: "extreme wide shot", "establishing wide", "vista wide", "environmental wide".
Stack the prompt in this order — each layer doing one job:
1. Shot type + ratio + lens. "Extreme wide shot, 2.39:1 anamorphic, 21mm lens." This is the load-bearing line; everything else modifies it.
2. Subject placement and scale. Ultra-wide collapses if you center the subject — name a third: "figure in lower-right third, occupying 15% of frame height". Small subject inside a large frame is what reads as scale.
3. Depth in planes, not bokeh. Ultra-wide lenses have deep depth of field by physics — asking for "shallow DOF" produces a fake, telephoto-looking image. Build depth with layered planes instead: "sharp textured foreground element 0–3ft, subject plane 8–15ft, midground architecture 30ft, atmospheric haze on distant ridge". Four planes is the working structure invideo's creative team uses for editorial sets — set design does the depth work, not the lens.
4. Camera motion. Name it precisely: "slow dolly-in, locked horizon", "lateral tracking left, parallax across foreground", "static wide, no camera move". Vague terms ("epic move") produce wobble.
5. Lighting + atmosphere. A single hard directional key with raked side light separates planes; volumetric haze or god rays read depth. "Single hard warm key, side-raked at 45°, atmospheric haze in midground, sky falls into shadow."
6. Negative prompt line. Close with what to exclude — this is where most ultra-wide prompts fail: "no centered subject, no portrait crop, no telephoto compression, no shallow depth of field, no fisheye distortion, no vertical framing."
For video specifically, model choice matters and the invideo agent routes per shot — it holds Runway, Veo, Kling, and Seedance 2.0, so you don't pick a platform per model. Seedance 2.0 handles multi-shot single-pass generation well, which preserves your wide framing across cuts; Kling 3.0 holds character scale inside wide environments; Veo reads cinematic lens vocabulary cleanly. State the model preference at the top of the prompt or let the invideo agent pick. For aspect ratio specifically, write it as a parameter line — "aspect: 2.39:1" or "--ar 21:9" — separate from the descriptive prose so the model parses it as a setting, not a suggestion.
One discipline that saves credits: lock the frame in an image first. Generate the still in your target ratio with Nano Banana Pro (best for lighting) or GPT-Image-2 (best for sharp environment detail), iterate the composition cheaply until the planes and placement hold, then pass that locked frame as a keyframe to Seedance 2.0 reference-to-video for the motion. Per invideo's creative director Hridaye, "One clean image works as the anchor for every future gen." Spend video credits only on a locked frame — across documented productions, image-first iteration is the primary mechanism keeping per-shot cost down (one editorial campaign produced 40 stills and 30 motion clips for ~630 credits / ~$150 total).
If the model still produces a flat or centered result, the fix is almost always one of three things: the ratio sits inside the descriptive prose instead of as a parameter; the subject has no explicit thirds placement; or you asked for shallow DOF on a wide lens. Move the ratio out, place the subject by thirds, and replace bokeh language with layered foreground/midground/background elements.
Adjacent: for shot consistency across an ultra-wide sequence, lock the first shot of the setup completely before generating the rest — subsequent shots in the same setup inherit the look and the framing logic.
Watch some of these to see what works for you:
One clean image works as the anchor for every future gen.
— Hridaye, invideo's creative director