VEO 3.1

Create promos, social media videos, explainers, product demos, real estate tours, and more with Google’s latest AI video model Veo 3.1. Invideo agent handles the prompting; you just direct.

Trusted by teams at
Phantom X
Wonder Studios
Google
NVIDIA
Salesforce
Visa
IBM
monday.com
Cisco
Mastercard
Meta
LinkedIn
Netflix
Ford
Hilton
Siemens
HubSpot
Phantom X
Wonder Studios
Google
NVIDIA
Salesforce
Visa
IBM
monday.com
Cisco
Mastercard
Meta
LinkedIn
Netflix
Ford
Hilton
Siemens
HubSpot

Why serious creatives choose Veo 3.1

Native audio generation

Sound generates in the same pass as the picture: dialogue between characters, effects tied to the action, ambience that matches the room. Cue a line in quotes and it arrives spoken in the shot; describe tires screeching and the screech lands where the tires do.

Ingredients to video

Give Veo up to three reference images: a person, a character, a product. It builds the shot around them and holds their identity in the output, so the face in your reference is the face on screen, take after take.

First and last frame

Give it a first frame and a last frame and it generates the shot between them: the motion, the camera path, and the sound that connect the two stills. The shot begins and ends exactly where you decided.

Scene extension

A scene continues off its own final second, seven seconds per pass, into a single continuous take of up to 148 seconds. Action, camera, and audio carry across every join, so the extension reads as one shot rather than a cut.

Prompt adherence

Cinematic language reads as instruction: dolly in, tracking shot, wide establishing, shallow focus, warm tones, and the output follows it as written. The shot you described is the shot you get, down to the lens.

Up to 4K, landscape or vertical

Generations run at up to 4K, whether the shot is headed for a screening or a feed. Vertical composes for the 9:16 frame from the first pass, so nothing arrives cropped to fit.


How Veo 3.1 works with invideo agents

On invideo, Veo 3.1 runs inside an agentic workflow: the agent picks it for the shots it does best, writes its prompts and audio cues in Veo’s own syntax, attaches your reference images to every generation, and runs Veo’s scene extension when a take needs to run long, all organized for your approval. Here is what that looks like in practice.

The agent picks the right model intelligently.

Veo 3.1 sits on a roster of 200+ models, and the agent routes work to it where it wins: dialogue scenes, sound-led moments, and cinematic realism that has to read as footage. You never choose a model unless you want to, and if you have routing rules of your own, they hold for the whole project.

You direct, the agent writes the prompts.

Veo rewards structured prompts, subject, action, camera, composition, lens, ambience, with dialogue cued in quotes and sound effects described outright. You call the shot in plain language and the agent writes it in Veo’s syntax, audio cues included, so the sound you imagined arrives inside the shot.

Your references reach the model on every generation.

Approve a reference image for each character and product, and the agent attaches up to three with every Veo generation it runs, so the model holds that exact appearance and the face in take twelve is the face in take one.

Long takes extend natively.

When a scene needs more than eight seconds, the agent uses Veo’s scene extension, continuing the take off its final second, seven seconds at a time, up to 148 seconds of continuous footage with sound intact. You ask the scene to keep going; the agent manages the passes.

You always stay in control.

You set how much the agent does on its own: let it generate everything, ask before videos, or see every prompt before it generates. That setting is yours to make, and yours to change.


Who is Veo 3.1 for?

Veo 3.1 supports cinematic, sound-complete video generation across filmmaking, marketing, and production workflows. With invideo agents underneath it, each team gets the model without learning the machinery.

Filmmakers

Create dialogue scenes, cinematic pre-visualization, first-and-last-frame shots, and sound-led motion tests with audio generated inside the clip. Veo 3.1 helps performances, camera moves, and room tone arrive together.

Microdrama producers

Build vertical episodes with native dialogue, reference-held characters, and scene extension when a beat needs to keep going. Useful for social-first stories where sound, realism, and continuity have to survive across episodes.

Performance Ads and UGC teams

Create talking-head UGC, product explainers, testimonial-style clips, and social ads with spoken dialogue, ambience, and effects already synced. Veo 3.1 is useful when the ad needs to read like footage, not a silent generation.

Brand and product marketers

Generate product films, demos, promos, real estate tours, and explainer videos in up to 4K, across landscape or vertical formats. Useful when the shot needs to open on the product, close on the logo, and sound finished from the first pass.

Helping creatives stay creative

Multiplayer mode

Collaborate in real time with live cursors to show what everyone's working on.

RRebeccajust now
Can we push the train arrival 2 sec later? Feels rushed.
AAiko2m
Love the warmth here — keep this lighting for the reunion shot.

Storyboarding

Turn any script or idea into a shot-by-shot plan, then tweak as needed before generating.

Script writing

Write your script inside invideo, and ask an AI co-writer for help if you'd like.

Timeline editor

Picture Premiere Pro with full AI.

Build your own agents

Create custom agents to fill specific roles like cinematographer, music designer, and more.

From solo creatives to creative enterprises

Private & SecureSOC 2 & ISO AlignedGDPR Compliant

World-class investors stand behind invideo.

Backed by the firms behind Stripe, Spotify, Flipkart, and ByteDance.

Pricing

All paid plans include:

Access to 200+ image, video, audio, music models including Seedance 2.0, Veo 3.1, Kling 3.0, Nano banana pro & Elevenlabs music.

Access to top stock providers like iStock, Storyblocks & more.

Model & agent prices are subject to change.

On-demand credit top-ups available.

Google Veo 3.1 FAQs

What is VEO 3.1?

Veo 3.1 is Google DeepMind’s latest AI video model, generating cinematic video with native audio, dialogue, sound effects, and ambience created in the same pass as the picture, at up to 4K, in landscape or vertical. On invideo, it runs inside the agent’s roster: you direct in plain language and the agent writes the instructions.

How can you use VEO 3.1 in invideo?

Three ways: add “use VEO 3.1” to your prompt, create a project through the Agents & Models flow, or pick “Clip from model” under Plugins and choose Veo 3.1. In every flow, the agent writes the model-facing prompt for you.

What kind of videos can I create using VEO 3.1 in invideo?

Anything built from cinematic shots with sound: product films, UGC-style ads, brand promos, real estate tours, explainers, and narrative scenes. The showcase on this page is all Veo work, from a Korean skincare product film to a retro car ad.

How long does it take to generate a video using VEO 3.1?

Typically from under a minute to a few minutes per clip, depending on resolution and load.

Do I need to learn VEO 3.1’s prompting to use it on invideo?

No. The agent writes to Veo in its own syntax, camera, composition, and audio cues included. You brief the shot the way you would brief a crew, and the agent handles the model’s language.