Blog

FLUX 3: Black Forest Labs' Video Model With Native Audio, Explained (August 2026)

Last updated August 7, 2026

FLUX 3: Black Forest Labs' Video Model With Native Audio, Explained (August 2026)

FLUX 3 is Black Forest Labs' multimodal model announced July 23, 2026: a single Self-Flow backbone generating images, up to 20-second ~720p video with natively synced audio, and robot actions. BFL's own evals show 93% wins over Ray 3.2 but ~52% parity with Seedance 2.0 — no independent data yet. Early Access only, no public pricing; it appears in the invideo agent's video roster.

Updated August 2026

FLUX 3 is Black Forest Labs' first video model — announced July 23, 2026 — and it is built differently from every other entry in the current video race. Instead of a dedicated video network, FLUX 3 is a single unified backbone that generates images, video, synchronized audio, and even robot-action sequences from one set of weights, under what the lab calls its "Real World Models" program. Concretely for creators: up to 20 seconds of ~720p video in a single generation, with natively synced sound, from text, images, or existing video. It is in Early Access as of August 2026, and it already appears in the invideo agent's video model roster.

What exactly is FLUX 3?

FLUX 3 is a multimodal generative model, not a video-only one. Black Forest Labs scaled it up from a research checkpoint published in March 2026, using an architecture the lab calls Self-Flow. The consequential design choice is the shared backbone: the same model that renders a frame also continues the motion, generates the soundtrack, and — in its robotics configuration — outputs physical action commands. The lab's framing, per its July 2026 announcement, is that video, audio, and action are all views of the same underlying world model rather than separate products.

For anyone tracking the company: this is the same ~70-person team behind the FLUX.1 and FLUX.2 image families, applying the "open weights eventually, API first" playbook it has used since 2024 — an open FLUX 3 Dev checkpoint is promised for later in 2026.

What can FLUX 3 generate?

Capability Detail (as of Aug 2026, per BFL's announcement)
Clip length Up to 20 seconds in a single generation
Resolution ~720p
Audio Native, synchronized — dialogue, ambience, and effects generated with the video
Input modes Text-to-video, image-to-video, video-to-video
Continuity tools Keyframe conditioning and clip continuation
Dialogue Multilingual spoken dialogue
Typography Animated on-screen text

Two of these are genuinely uncommon. A 20-second single generation is roughly double the typical 8–10-second ceiling most current video models work within, which matters for dialogue scenes and continuous camera moves that stitching can't fake cleanly. And video-to-video plus continuation means FLUX 3 can extend or restyle existing footage, not just start from scratch.

Why is FLUX 3's native audio a big deal?

Because of an economics argument, not merely a features one. Black Forest Labs disclosed that audio accounts for less than 0.5% of FLUX 3's training tokens. Sound is simply far less information-dense than pixels, so once you commit to a unified multimodal backbone, adding synchronized audio costs almost nothing in compute — the video training dominates the bill either way.

The implication cuts across the industry: models that generate silent video and bolt sound on later are paying an integration cost for something that is nearly free inside a unified architecture. Synced dialogue and effects stop being a premium add-on and become a default property of the model. That is a structural bet, and FLUX 3 is the clearest test of it so far — worth watching whether the rest of the field converges on the same design through 2026–27.

How does FLUX 3 compare with other video models?

Caveat first: every number in this section is Black Forest Labs' own evaluation. These are first-party win rates from the July 2026 announcement, measured on 10-second 720p generations, and no independent arena data exists yet as of August 2026. Treat them as the lab's claim, not settled ranking:

Matchup (BFL's own evals, 10s/720p) FLUX 3 win rate
vs Ray 3.2 93%
vs Gen-4.5 77%
vs Grok Imagine 69%
vs Kling v3 Pro 60%
vs Seedance 2.0 ~52%
vs Gemini Omni Flash ~52%

Read honestly, the lab's own numbers say FLUX 3 clearly beats the mid-field but sits at rough parity with the strongest current models — Seedance 2.0 and Gemini Omni Flash. A lab publishing a ~52% self-reported result against the leaders is at least not pretending to dominance, but until FLUX 3 shows up in blind arenas, the fair summary is: credibly frontier-adjacent, unverified.

What does the robotics work say about where this is going?

The same backbone powers FLUX-mimic, Black Forest Labs' robot-control configuration, which the lab says reacts to visual input in 101 milliseconds and adapts to new tasks with roughly 30-minute fine-tunes — and is already running in a production deployment at Audi (as of the July 2026 announcement). For video creators this matters as evidence, not as a feature: a model that can drive a physical robot arm from the same weights that render a clip is being trained to predict how the world actually behaves. That is the strongest available signal that FLUX 3's video physics — object permanence, contact, momentum — come from world-model ambitions rather than pure visual mimicry.

When can you actually use FLUX 3, and what does it cost?

As of August 2026, FLUX 3 is Early Access only. Video and action generation are live for early partners; image generation on the FLUX 3 backbone is slated to follow "in the coming weeks," per the announcement. A public API and the open FLUX 3 Dev weights are promised for later in 2026. There is no public pricing — any per-second or per-clip price you see quoted today is a reseller's invention, not an official number.

What people ask about FLUX 3

Is FLUX 3 available to the public?

Not yet. As of August 2026 it is in Early Access with selected partners, with a general API and open FLUX 3 Dev weights promised later in 2026. It is, however, already selectable inside the invideo agent's video roster.

Does FLUX 3 generate sound?

Yes — natively. Dialogue (multilingual), ambience, and effects are generated in sync with the video by the same model, not layered on afterward. BFL notes audio made up under 0.5% of training tokens, which is why native audio ships as a default rather than a paid add-on.

How long can a FLUX 3 video be?

Up to 20 seconds in a single generation at ~720p, with keyframe and continuation modes for building longer sequences from consistent segments.

Is FLUX 3 better than Seedance 2.0 or Kling v3 Pro?

Unknown — independently. Black Forest Labs' own evals report 60% wins over Kling v3 Pro and ~52% against Seedance 2.0 on 10-second 720p clips, i.e. rough parity with the strongest models by the lab's own measure. No third-party arena results exist yet as of August 2026.

Will FLUX 3 have open weights?

Black Forest Labs has promised an open "FLUX 3 Dev" checkpoint later in 2026, consistent with its FLUX.1 and FLUX.2 release pattern. License terms are unannounced — given the FLUX.2 precedent, check whether "open" means Apache 2.0 or the outputs-only non-commercial license before building on it.

How much does FLUX 3 cost?

There is no official pricing as of August 2026. Until BFL publishes rates, no quoted price is trustworthy.

Where can you try FLUX 3?

FLUX 3 is listed in the invideo agent's video model roster as "FLUX 3 — text, image, and video to video with audio by Black Forest Labs," sitting alongside Veo 3.1, Sora 2, Kling 3.0, Seedance 2.0 and the rest of the 200+ models invideo runs — which makes the agent a practical way to put BFL's first-party claims to the test against the models it was benchmarked on, per shot, in one timeline. Start from invideo.io, or go straight to a workflow: the AI video generator for text-to-video, image to video for animating stills with synced sound, or the AI animation generator for stylized motion.


Version history

  • August 2026 — First published, covering the July 23, 2026 announcement: Self-Flow unified backbone, 20s/720p native-audio video, first-party win rates, FLUX-mimic robotics, and Early Access status.
Share