Black Forest Labs Launches FLUX 3 Multimodal Foundation Model
Original: Flux 3
Why This Matters
Unified video-image-audio generation in one model marks a significant architectural shift in generative AI for physical and digital environments.
Black Forest Labs released FLUX 3 on July 23, 2026, a multimodal foundation model jointly trained on images, video, and audio within a unified architecture. Built on the company's Self-Flow framework, FLUX 3 is now available in Early Access via the bfl.ai platform.
Black Forest Labs announced FLUX 3, a new multimodal foundation model that jointly learns from images, video, and audio in a single unified architecture. The company describes the approach as building a 'representation of the world' rather than optimizing for any single modality, arguing that each data type — images, video, audio, and language — captures a different projection of the same underlying physical reality.
FLUX 3 is built on Self-Flow, Black Forest Labs' proprietary method for aligning multimodal generation and understanding within the same model. The company reports that Self-Flow outperforms standard Flow Matching (FM) on both generation error (measured by Fréchet distance) and manipulation task success rates after fine-tuning.
Key capabilities include: text-to-video and image-to-video generation with native audio up to 20 seconds at 720p; video-to-video generation preserving characters or scenes; generative video-audio continuation; keyframe-to-video transitions; multilingual dialogue; strong typography generation; and agentic chaining of clips into multi-shot sequences. All video outputs include native audio.
Black Forest Labs positions FLUX 3 as a step toward 'real-world visual intelligence' with early applications in content creation and physical AI. The model is currently in Early Access, and the company notes evaluations are preliminary with further improvements expected.