Jul 23, 2026
AI

Flux 3 video model adds native audio for clips up to 20 seconds

Black Forest Labs says Flux 3 generates 20-second videos with native audio and is testing a related robotics model at Audi.

Wei-Lin Zhao

By Wei-Lin Zhao · AI Correspondent

· 3 min read

Flux 3 video model adds native audio for clips up to 20 seconds
Photo: The Decoder

Black Forest Labs has released Flux 3, a multimodal foundation model whose Flux 3 video capabilities include generating clips of up to 20 seconds with native audio. The German AI company says the system trains on images, video and audio together, a design aimed at improving video generation and extending the model into robotics-related action prediction.

The release puts BFL into the same commercial and technical race as Runway, Luma, Google, ByteDance and other video-generation vendors. The company did not disclose model size, training costs, customer pricing or enterprise adoption, so the near-term significance rests on claimed capability gains and the planned rollout of access.

What can Flux 3 video do?

According to Black Forest Labs, Flux 3 supports text-to-video, image-to-video and video-to-video generation, along with keyframe-based transitions. The company also says it can produce multilingual dialogue and connect separate clips through agent-driven workflows to create longer multi-shot sequences.

Native audio is the notable addition. BFL says the model can generate sound alongside the video rather than treating audio as a separate post-production layer. The company claims the model is particularly strong at rendering human facial expressions and aligning sounds with physical events, although those claims have not yet been independently tested.

BFL frames Flux 3 as part of a broader move toward models that can represent physical and digital environments. Its argument is that images, video and audio each expose different signals: images capture spatial detail, video captures motion over time, and audio can indicate cause-and-effect relationships such as mechanical movement and sound. Training across all three is meant to give the model a shared internal representation rather than isolated media skills.

How did Flux 3 compare with rival video models?

Black Forest Labs published early preference results based on 10-second, 720p clips. The company said Flux 3 was favored against several competing systems, with the widest margins against Luma Ray 3.2 and Runway Gen-4.5. BFL described the results as preliminary, and no independent benchmark results are available yet.

  • 93% preference rate versus Luma Ray 3.2, according to BFL
  • 77% versus Runway Gen-4.5
  • 69% versus Grok Imagine Video
  • 60% versus Kling v3 Pro
  • 59% versus Happy Horse v1
  • 57% versus Happy Horse 1.1
  • 52% versus Seedance 2.0 and 52% versus Gemini Omni Flash

The closer results against Seedance 2.0 and Gemini Omni Flash matter because a 52% preference rate is near parity in this kind of comparison. BFL’s data suggests Flux 3 may be competitive with leading video systems, but company-run preference tests are not a substitute for broad external evaluation across prompt types, safety behavior, consistency and cost.

What is Flux-mimic and how is Audi involved?

BFL also introduced Flux-mimic, a video-action model developed with Mimic Robotics. The company says Flux-mimic is being tested by Audi on production tasks, making robotics the most concrete industrial use case disclosed with the Flux 3 launch.

Flux 3 is built on BFL’s Self-Flow approach, which the company describes as a way to train one model to both understand and generate content. The architecture uses a multimodal transformer with separate encoders and decoders for images, video, audio and actions, converting those inputs into a shared representation and then back into outputs.

BFL says the action component can be extended for robotics uses and that Self-Flow outperforms standard flow-matching methods in generation quality and physical-world understanding. Those are company claims, and BFL has not provided independent validation for the broader physical-intelligence assertions.

When will Black Forest Labs release Flux 3 access?

BFL says Flux 3 Video is available now, with Flux 3 Image planned for early access in the coming weeks. The company expects Flux 3 Image to improve image generation, especially for complex prompts and rendering text accurately across multiple languages.

Action prediction will first be available through selected partners. BFL also plans to release open-weight access to the multimodal backbone under the name Flux 3 Dev. Longer term, the company says it is working on models that combine perception, action and language prediction in a single system.

This story draws on original reporting from The Decoder.

More from AI

All AI →