Black Forest Labs has released FLUX 3, a multimodal model capable of generating synchronized audio and video output natively, with clips up to 20 seconds long1,2.
The model represents Black Forest Labs' first multimodal release, combining video generation with native audio generation in a single pipeline. FLUX 3 is described as the first model to support 20-second native audio-visual synchronization generation.
Latent Space characterized the FLUX 3 launch as more consequential than same-day releases from other labs, noting that OpenAI launched ChatGPT Voice for consumers and OpenAI Presence for enterprise on the same day, while the timing also coincided with Claude Voice3.
The Latent Space coverage references Black Forest Labs' earlier work on image generation, noting that the company initially launched with models hinting at video capabilities to come.
ANALYSIS The addition of native audio to video generation — rather than requiring a separate audio model — distinguishes the architectural approach from workflows that layer audio post-generation.
The 20-second synchronized audio-video generation capability marks a concrete expansion of Black Forest Labs' product surface from its earlier image-generation models into multimodal territory.