Gemini 3.1 Flash Text to Speech

Generate speech that performs the way you direct it, from a whisper to a shout, one narrator or a whole conversation. Describe the read to invideo’s agent, and it writes the direction Gemini 3.1 Flash TTS performs in.

Trusted by teams at
Phantom X
Wonder Studios
Google
NVIDIA
Salesforce
Visa
IBM
monday.com
Cisco
Mastercard
Meta
LinkedIn
Netflix
Ford
Hilton
Siemens
HubSpot
Phantom X
Wonder Studios
Google
NVIDIA
Salesforce
Visa
IBM
monday.com
Cisco
Mastercard
Meta
LinkedIn
Netflix
Ford
Hilton
Siemens
HubSpot

Why serious creatives choose Gemini 3.1 Flash TTS

It performs the direction, not just the words

Write the direction into the script, calm here, a whisper on this line, a shout on that one, and Gemini 3.1 Flash TTS performs it, with granular control over style, pace, tone, and delivery. The note you write is the read you get.

It can carry a whole conversation

From a single script, Gemini 3.1 Flash TTS generates multi-speaker audio: two characters arguing, a podcast exchange, an interview, a dramatic scene, each voice distinct, the timing of a real conversation. Dialogue stops being one voice at a time.

It paces like a person, not a teleprompter

The rushed, even cadence that gives TTS away is not here. Gemini 3.1 Flash TTS breathes where a speaker would, lands emphasis where the sentence means it, and holds a natural rhythm from short snippets to long-form speeches.

It sounds native in every language

In every language it speaks, Gemini 3.1 Flash TTS keeps the local rhythm: stress where that language puts it, pacing the way its speakers actually talk, not an accent pasted over English cadence.

It holds up at length

From a one-line snippet to a long-form speech, the performance stays consistent, the same voice, the same energy, no drift as the read goes on.


How Gemini 3.1 Flash TTS works with invideo agents

On invideo, Gemini 3.1 Flash TTS runs inside an agentic workflow: you direct in plain language, and the agent turns that direction into the performance instructions the model follows. Here is what that looks like in practice.

Your direction becomes the performance.

Say the apology should start defensive and end sincere, and the agent writes that arc into the script as the audio direction Gemini 3.1 Flash TTS performs, line by line, shift by shift. You direct like a director; the agent handles the notation.

The agent directs the whole exchange, not one line at a time.

Because the agent holds your full script in its memory, it knows who is in the scene and who says what: it casts the voices, assigns the dialogue, and has Gemini 3.1 Flash TTS perform the full exchange in one pass, with the interplay of a real conversation rather than two reads spliced together.

Voices stay cast, scene after scene.

The voices you approve stay bound to their characters in the project’s memory, so the same speaker carries their part across every scene, and a note to one voice never touches the other.

You always stay in control.

How much the agent runs on its own is your setting, not its default: full takes delivered for review, or every direction shown before it generates. The read that ships is the one you signed off on.


Who is Gemini 3.1 Flash TTS for?

AI filmmakers and microdrama producers.

Dialogue is where audio work gets hard. Gemini 3.1 Flash TTS performs the scene as written, two characters, real interplay, and takes direction on every line. Useful whether you are directing AI filmmaking or a microdrama season.

Animators.

Every character needs a voice, and Gemini 3.1 Flash TTS provides the cast: performances directed line by line, dialogue scenes delivered as real exchanges, and a retake never further than a note.

Explainer and educator makers.

Gemini 3.1 Flash TTS narrates a full lesson the way a good teacher delivers it: emphasis where the point lands, a shift in tone when the topic turns, and pacing that holds attention to the end.

Helping creatives stay creative

Multiplayer mode

Collaborate in real time with live cursors to show what everyone's working on.

RRebeccajust now
Can we push the train arrival 2 sec later? Feels rushed.
AAiko2m
Love the warmth here — keep this lighting for the reunion shot.

Storyboarding

Turn any script or idea into a shot-by-shot plan, then tweak as needed before generating.

Script writing

Write your script inside invideo, and ask an AI co-writer for help if you'd like.

Timeline editor

Picture Premiere Pro with full AI.

Build your own agents

Create custom agents to fill specific roles like cinematographer, music designer, and more.

From solo creatives to creative enterprises

Private & SecureSOC 2 & ISO AlignedGDPR Compliant

World-class investors stand behind invideo.

Backed by the firms behind Stripe, Spotify, Flipkart, and ByteDance.

Pricing

All paid plans include:

Access to 200+ image, video, audio, music models including Seedance 2.5, Veo 3.1, Kling 3.0, Nano banana pro & Elevenlabs music.

Access to top stock providers like iStock, Storyblocks & more.

Model & agent prices are subject to change.

On-demand credit top-ups available.

Gemini 3.1 Flash TTS FAQs

What is Gemini 3.1 Flash TTS?

Gemini 3.1 Flash TTS is Google's text to speech model, and on invideo agent it is the one that performs a direction rather than just reads a line. You describe how the read should go and the agent writes the performance instructions the model follows.

Can I control how the voice delivers a line?

Yes, line by line, and invideo agent is how you do it. Calm here, a whisper on this line, a shout on that one, an apology that starts defensive and ends sincere: you say it in plain language and the agent turns it into the direction Gemini 3.1 Flash TTS performs, with control over style, pace, tone, and delivery.

Can it generate a conversation between two speakers?

Yes. Because invideo agent holds your full script in memory, it knows who is in the scene and who says what, so it casts the voices, assigns the dialogue, and has Gemini 3.1 Flash TTS perform the exchange in one pass, with the interplay of a real conversation rather than two reads spliced together.

Does it work in other languages?

Yes, and invideo agent will direct the read in any of them. Gemini 3.1 Flash TTS keeps the local rhythm in each language, stress where that language puts it and pacing the way its speakers actually talk, instead of an accent pasted over English cadence.

Do I need to learn audio tags to direct it on invideo?

No. You describe the performance to invideo agent the way you would brief a voice actor, and it writes the notation Gemini 3.1 Flash TTS reads. There are no tags to learn and no syntax to get right, and every take comes back for your approval.