

Arooj Ishtiaq
July 27, 2026 • Updated July 27, 2026
13 mins Read
Black Forest Labs announced FLUX 3 on July 23, 2026. It's a genuinely different kind of model from what came before it.
One architecture now trains jointly on images, video, and audio. That replaces the old approach of stitching three separate systems together behind a shared interface.
This guide covers every FLUX 3 feature confirmed so far. It breaks down what's actually available today versus what's still rolling out, and how the model stacks up against the video generators creators already use.
FLUX 3 Features at a Glance
| Specification | Details |
|---|---|
| Announced | July 23, 2026 |
| Built By | Black Forest Labs (founded by researchers behind Stable Diffusion and latent diffusion models). |
| Core Architecture | Self-Flow architecture trained jointly across images, video, and audio for unified multimodal generation. |
| Maximum Video Length | Up to 20 seconds per clip, with support for chaining clips into multi-minute sequences. |
| Video Resolution | 720p during the early access period. |
| Native Audio | Yes. Audio is generated in the same inference pass as the video for synchronized output. |
| Available Now | FLUX 3 Video and FLUX 3 Action (both available through gated early access). |
| Not Yet Launched | FLUX 3 Image and FLUX 3 Dev (open-weight version). |
| Public API | Not yet available for any FLUX 3 tier. |
What Is FLUX 3?
FLUX 3 is Black Forest Labs' first multimodal foundation model. It's built on an architecture called Self-Flow that aligns generation and understanding across images, video, and audio simultaneously, rather than training a separate model for each.
The reasoning behind that design is worth understanding. It explains why FLUX 3 behaves differently from a typical video generator.
No single modality fully describes reality:
- Images capture spatial structure at one instant.
- Video restores the dimension of time and reveals how things move.
- Audio reveals the causal link between physical events and the sounds they produce, something vision alone can't detect.
Training across all three at once means each modality constrains the others. Sound has to match the impact on screen. Motion has to obey the mass of the object moving. A future frame has to follow logically from the past one. Black Forest Labs describes this as building one working representation of the world, rather than three disconnected projections of it.
That framing matters. It positions FLUX 3 as more than an image generator that learned to make video. It's a single model meant to eventually cover content generation, world simulation, and physical action prediction under one architecture.
The team behind it has real credibility for that ambition:
- Black Forest Labs was founded by the researchers who created latent diffusion, the technique underlying Stable Diffusion.
- The company's existing FLUX models already power generative features inside Adobe Photoshop, Picsart, and other production tools.
That track record is part of why FLUX 3's claims are worth taking seriously even before every feature ships.
FLUX 3 Video Features
FLUX 3 Video is the first and, at launch, the most broadly available part of the release.
FLUX 3 Video Capabilities
- Clips up to 20 seconds in a single generation, with resolution capped at 720p during early access.
- Native synchronized audio generated in the same pass as the video, matching the physical events on screen rather than applying a generic soundtrack afterward.
- Agentic clip chaining, which links individual clips into longer, multi-shot sequences lasting several minutes, using visual references to keep characters and style consistent across every cut.
- Wide style range, spanning candid camcorder-style footage, UGC-style clips, full cinematic output, and animation, all from one model without switching tools.
- Strong typography and animated design, rendering legible on-screen text and title sequences directly inside generated video.
FLUX 3 Video Generation Modes
FLUX 3 supports six distinct generation modes:
| Generation Mode | Description |
|---|---|
| Text-to-Video | Generates a complete video clip using only a written text prompt, without requiring any visual reference. |
| Image-to-Video | Animates a still image, continues from a starting frame, or uses an image as a visual style and composition reference. |
| Video-to-Video | Transforms an existing video while preserving key elements such as characters, objects, motion, or overall style in a new scene. |
| Keyframe-to-Video | Creates a controlled animation between two or more defined keyframes, producing smooth visual transitions. |
| Video + Audio Continuation | Extends an existing video while seamlessly continuing synchronized dialogue, music, sound effects, and ambient audio. |
| Multilingual Dialogue | Generates realistic character speech in multiple languages with synchronized lip-sync and natural facial animation. |
Early testers working through the initial access rollout have flagged one recurring issue: image references don't always attach to the output as consistently as text-to-video mode performs. This reads as a launch-stage rough edge rather than a fundamental limitation.
Once a clip generates the way you want it, a production platform's AI video editor is the natural next step for trimming and final polish.
FLUX 3 Image Features
FLUX 3 Image is not yet broadly available. As of this guide's publication, it remains in pre-launch evaluation. Black Forest Labs has stated it will enter early access in the weeks following the initial video and action rollout.
What the company has shown so far:
- Significantly improved handling of complex, multi-element prompts compared to earlier FLUX versions.
- High-accuracy text rendering across multiple languages inside generated images.
- A wide output range spanning illustration, photography, product renders, and fine art at flexible aspect ratios and resolutions.
Because these results come from Black Forest Labs' own mid-training evaluations rather than a public release, treat the image capabilities as a preview of direction rather than a confirmed, testable feature set. An overview of FLUX 2 and a look at FLUX Kontext cover the versions FLUX 3 Image is expected to build on.
FLUX 3 Action: Video Meets Robotics
The least conventional part of the release is FLUX 3 Action. It extends the same world-understanding architecture into physical action prediction for robotics.
Black Forest Labs took two routes to this:
- Building native action prediction directly into FLUX 3.
- Using the pretrained video backbone as a foundation that specialized robotics models can fine-tune from, with limited task-specific data.
The company's first partner on this front is Mimic Robotics. Together they built FLUX-mimic, a combined video-action model aimed at dexterous manipulation tasks. Audi is reportedly testing the system on a real production line, a notably concrete proof point compared to most AI robotics announcements.
The underlying argument, that video generation and physical action prediction are the same problem viewed from different angles, is the most ambitious claim in the entire FLUX 3 announcement. It's also the one that will take considerably longer to validate than the video or image features.
FLUX 3 Availability: What's Live and What Isn't
This is where a lot of early coverage gets ahead of what's actually accessible, so it's worth stating plainly.
| FLUX 3 Product | Status | Availability |
|---|---|---|
| FLUX 3 Video | Live | Available through gated early access. Users must submit an application to receive access. |
| FLUX 3 Action (FLUX-mimic) | Live | Available only to selected robotics companies and research partners during the early access phase. |
| FLUX 3 Image | Not Yet Launched | Expected to launch in the coming weeks. |
| FLUX 3 Dev (Open-Weight) | Not Yet Launched | Planned for release later in 2026. |
A few things worth knowing about accessing FLUX 3 right now:
- There is currently no public API access to any FLUX 3 tier from Black Forest Labs or its partners.
- No pricing has been announced for any tier.
- Named early testing partners include Canva, Burda, Magnific (formerly Freepik), Krea, and Picsart, giving Black Forest Labs several existing creative-tooling routes to put the model in front of real users quickly.
- This staged rollout mirrors the pattern several frontier labs have used recently for major model releases, though Black Forest Labs hasn't framed its own staging around the security rationale other companies have cited.
FLUX 3 vs. Other Video Models: Performance Comparison
Black Forest Labs published preference comparisons from its own early evaluations, generated using 10-second, 720p text-to-video clips with audio.
| AI Video Model | Arena Score |
|---|---|
| Luma Ray 3.2 | 93% |
| Runway Gen-4.5 | 77% |
| Grok Imagine Video | 69% |
| Kling v3 Pro | 60% |
| Happy Horse v1 | 59% |
| Happy Horse 1.1 | 57% |
| Seedance 2.0 | 52% |
| Gemini Omni Flash | 52% |
Two things are worth keeping in mind before treating this table as a settled ranking:
- Black Forest Labs itself labels the chart a "preliminary evaluation of an early FLUX 3 candidate," meaning the results reflect a pre-release checkpoint rather than necessarily the exact model now available in early access.
- The strongest margins, 93% against Luma Ray 3.2 and 77% against Runway Gen-4.5, come against comparison points that aren't currently setting the pace in independent video rankings.
The closer results tell the more useful story. A near coin-flip against Seedance 2.0 and Gemini Omni Flash is the more meaningful data point, and also the one Black Forest Labs has committed to revisiting with fuller methodology once the model reaches broader availability.
For side-by-side testing once access opens more broadly, a video generator hub covering models including Seedance 2.5, Kling 3.0, Veo 3.1, Gemini Omni Flash, and Hailuo 3.0 lets you compare the same prompt across models directly rather than relying on any single lab's internal benchmark.
FLUX 3 vs. FLUX 2: What Actually Changed
FLUX 2 was an image-focused model with strong realism and editing capability. FLUX 3 is a structural departure rather than an incremental upgrade.
Two things separate them:
- FLUX 3 is the first FLUX model trained jointly across video and audio rather than image alone.
- It extends the same underlying architecture toward physical action prediction, a category FLUX 2 never touched.
On the image side specifically, FLUX 3 claims meaningfully better complex-prompt handling and text rendering than FLUX 2. That comparison won't be independently testable until FLUX 3 Image reaches early access.
Anyone deciding whether to wait for FLUX 3 or keep working in FLUX 2 today should treat the video and action features as the genuinely new territory, and the image improvements as directionally promising but unverified until release.
Who Should Use FLUX 3
- Filmmakers and narrative creators. Agentic clip chaining and keyframe-to-video give directorial control over multi-shot sequences without manual compositing between clips. The model's physical accuracy in motion and audio keeps action sequences and dialogue coherent across a longer cut.
- Brand and campaign teams. High-accuracy text rendering inside generated video and images matters directly for ads, packaging visuals, and any branded creative where legible on-screen copy is part of the concept, not an afterthought added in post.
- Global and multilingual content teams. Native dialogue generation with accurate lip-sync across languages removes a step that traditionally required separate dubbing or localization work entirely.
For structuring prompts to get consistent results out of a multimodal model like this, a guide to JSON prompting for AI video and the equivalent guide for AI image prompting cover how structured prompts improve consistency across generations. That technique matters more, not less, as models take on more simultaneous inputs.
Conclusion
FLUX 3's real news isn't any single feature; it's the architecture underneath all of them. One model is learning images, video, and audio together well enough that the same backbone now extends into robotics.
Video is the part creators can actually test today. Image is close behind. The preference numbers Black Forest Labs has published so far are promising but explicitly preliminary.
A dedicated FLUX 3 workspace gives creators a way to try the model's video generation modes directly and compare output against other leading models in one place as access continues to expand.
Frequently Asked Questions
What is FLUX 3?
FLUX 3 is Black Forest Labs' first multimodal foundation model, trained jointly on images, video, and audio through an architecture called Self-Flow. It generates images, video clips up to 20 seconds with native synchronized audio, and extends toward physical action prediction for robotics, all from one underlying model.
Is FLUX 3 available to the public?
Partially. FLUX 3 Video and FLUX 3 Action are in gated early access, requiring an application that Black Forest Labs must approve. FLUX 3 Image has not launched yet and is expected in the coming weeks. FLUX 3 Dev, the open-weight version, is planned for later in 2026. No tier currently has public API access or announced pricing.
How long can FLUX 3 videos be?
Up to 20 seconds in a single generation pass, at 720p resolution during early access. Agentic clip chaining extends this into multi-minute sequences by linking multiple clips with consistent characters and style.
Does FLUX 3 generate audio automatically?
Yes. Native audio is generated in the same pass as the video by default, synchronized to the physical events on screen, including multilingual dialogue with lip-sync where the prompt calls for it.
How does FLUX 3 compare to Kling, Seedance, and Veo?
In Black Forest Labs' own early evaluations, FLUX 3 was preferred in roughly half of comparisons against Seedance 2.0 and Gemini Omni Flash, essentially a statistical tie, and by wider margins against Luma Ray 3.2 and Runway Gen-4.5. The company has labeled these results preliminary and plans to publish fuller benchmark methodology at broader release.
What is FLUX 3 Action?
FLUX 3 Action extends FLUX 3's world-understanding architecture into physical action prediction for robotics, developed with robotics partner Mimic under the name FLUX-mimic. It's aimed at dexterous manipulation tasks and is reportedly being tested by Audi on a production line, separate from FLUX 3's content-generation features.

Arooj Ishtiaq
Arooj is a SaaS content writer specializing in AI models and applied technology. At ImagineArt, she creates sharp, product-focused content that helps creators and businesses understand, adopt, and get real value from AI tools.