Black Forest Labs Launches FLUX 3 Multimodal Foundation Model

Original: Flux 3

Why This Matters

Unified video-image-audio generation in one model marks a significant architectural shift in generative AI for physical and digital environments.

Black Forest Labs released FLUX 3 on July 23, 2026, a multimodal foundation model jointly trained on images, video, and audio within a unified architecture. Built on the company's Self-Flow framework, FLUX 3 is now available in Early Access via the bfl.ai platform.

Black Forest Labs announced FLUX 3, a new multimodal foundation model that jointly learns from images, video, and audio in a single unified architecture. The company describes the approach as building a 'representation of the world' rather than optimizing for any single modality, arguing that each data type — images, video, audio, and language — captures a different projection of the same underlying physical reality.

FLUX 3 is built on Self-Flow, Black Forest Labs' proprietary method for aligning multimodal generation and understanding within the same model. The company reports that Self-Flow outperforms standard Flow Matching (FM) on both generation error (measured by Fréchet distance) and manipulation task success rates after fine-tuning.

Key capabilities include: text-to-video and image-to-video generation with native audio up to 20 seconds at 720p; video-to-video generation preserving characters or scenes; generative video-audio continuation; keyframe-to-video transitions; multilingual dialogue; strong typography generation; and agentic chaining of clips into multi-shot sequences. All video outputs include native audio.

Black Forest Labs positions FLUX 3 as a step toward 'real-world visual intelligence' with early applications in content creation and physical AI. The model is currently in Early Access, and the company notes evaluations are preliminary with further improvements expected.

Source

bfl.ai — Read original →