Black Forest Labs' Flux 3 Adds Native Audio to AI Video

LLMsAI Agents
Illustration generated by AI for this article

The Core · TL;DR

  • Black Forest Labs launched Flux 3, its first model to generate video with native, synchronized audio, in clips up to 20 seconds long.
  • Built on BFL's Self-Flow architecture, Flux 3 uses a multimodal transformer with dedicated encoders/decoders for images, video, audio, and robotic actions.
  • BFL's internal benchmarks claim Flux 3 beats Luma Ray 3.2 (93%), Runway Gen-4.5 (77%), and other rivals, though these results are unverified by third parties.
  • A robotics variant, Flux-mimic, built with Mimic Robotics, is being tested on manufacturing tasks at Audi; Flux 3 Video pricing ranges from $0.06 to $0.53 per second depending on mode and quality.

Black Forest Labs has released Flux 3, a multimodal foundation model that generates video with synchronized audio baked in from the start, rather than layered on afterward. The company says it's the first time its Flux line has produced clips with native sound, and the resulting videos can run up to 20 seconds long.

That native-audio capability is the headline feature. Flux 3 Video, now in early access, can render lip-synced dialogue in more than 14 languages along with sound effects and ambient noise, all generated in the same pass as the visuals rather than added through post-processing.

The model is built on Self-Flow, BFL's architecture for training a single system to both generate and interpret content simultaneously. Under the hood, Flux 3 uses a multimodal transformer with separate encoders and decoders handling images, video, audio, and, notably, robotic actions.

That last piece points to Flux 3's ambitions beyond content creation. BFL partnered with Mimic Robotics to build Flux-mimic, a video-action variant now being trialed on real manufacturing tasks at Audi, aimed at robot manipulation in factory settings.

Capabilities and early benchmarks

Beyond text-to-video and image-to-video, Flux 3 supports video-to-video transformation, keyframe-based scene transitions, multilingual dialogue generation, and agent-driven stitching of clips into longer multi-shot sequences.

BFL published internal comparison results using 10-second, 720p clips. Flux 3 reportedly won 93 percent of head-to-head comparisons against Luma Ray 3.2, 77 percent against Runway Gen-4.5, and 69 percent against Grok Imagine Video.

Margins were tighter against other competitors: 60 percent against Kling v3 Pro, roughly 57-59 percent against Happy Horse's two versions, and 52 percent apiece against Seedance 2.0 and Gemini Omni Flash. BFL also cites Elo scores of 1,135 for text-to-video and 1,051 for image-to-video.

BFL itself cautions that these evaluations are internal and preliminary, with no independent testing yet available to confirm the claims.

Pricing and availability

Flux 3 Video is priced by mode and resolution. Draft-mode generation costs $0.06 per second for text-to-video or image-to-video and $0.12 per second for video-to-video.

Full-quality HD pricing rises to $0.17 and $0.41 per second respectively, while Full HD reaches $0.29 and $0.53 per second. The model is accessible through BFL's API and select partners.

Flux 3 Image, the still-image counterpart, is not yet available but BFL says it plans to bring it to early access within the coming weeks. For now, video generation with synchronized sound is the model's clearest differentiator against rivals racing to combine visual and audio generation in one system.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram