
Hey guys, Mr. Technology here.
It is Tuesday, July 28, 2026, and the most architecturally interesting release of last week was FLUX 3 from Black Forest Labs — the Freiburg-based lab that invented Stable Diffusion and built every FLUX release for three years.
BFL did the thing everyone has been talking about and nobody has shipped. FLUX 3 is one model, trained jointly across image, video, audio, and action-prediction, that generates 20-second video with native audio in a single pass. Not a router. Not a pipeline of separate models behind a common API. One set of weights, one architecture, one forward pass.
What You Need to Know
>
- Black Forest Labs released FLUX 3 in Early Access on July 23, 2026 as a multimodal foundation model jointly trained on image, video, and audio using their Self-Flow approach. - It generates videos up to 20 seconds with native audio in a single pass, plus image-to-video, video-to-video, keyframe-to-video, multilingual dialogue, and action prediction for robotics. - Four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action (gated Early Access now), and FLUX 3 Dev (open weights, later this year). - BFL's preliminary evaluations put it ahead of every major video model: 77% over Runway Gen-4.5, 93% over Luma Ray 3.2, 69% over Grok Imagine Video, 60% over Kling v3 Pro. - No pricing, no public API, no downloadable weights at launch.
The 2026 release cycle has been dominated by frontier text labs. OpenAI shipped GPT-5.5 and the GPT-5.6 family. Anthropic finished Claude 5 with Opus 5 on July 24. Google dropped Gemini 3.6 Flash. DeepSeek pushed V4 to GA. xAI shipped Grok STT 1.0. Every headline was a text model with vision attached. FLUX 3 is a different category. It is the first model from a credible frontier lab that was trained from scratch on three modalities simultaneously and treats audio as a first-class output rather than an afterthought bolted on by a separate TTS pass. The Self-Flow approach BFL published alongside the release is the technical core: flow matching across modalities, with the modalities acting as mutual constraints during training instead of being aligned post-hoc.
The practical consequence: the audio and video come out temporally aligned by construction. The dog barks at the moment the dog appears. The footsteps match the footfall. If you have ever stitched together a separate video model and a separate audio model and tried to make them agree, you know how hard this is to fake after the fact.
Four product lines:
BFL's preliminary analysis used 10-second text-to-video clips in 720p with audio. Results are expected to improve during Early Access.
BFL is running internal A/B comparisons against the current video-model lineup. Their published preferences for FLUX 3 against alternative outputs for the same prompts:
Two honest caveats. First, BFL has not published sample sizes, rater counts, or the full evaluation methodology — these are self-reported preference rates, and VentureBeat and Latent Space both flagged the missing detail. Second, the comparisons are against video competitors, and the joint audio generation is not scored separately because none of the competing video models emit native audio at all.
Read the numbers as "FLUX 3 wins the video-preference sweep in BFL's own setup." Still meaningful — Runway and Kling do not lose to many models — but it is not a frontier-text-model benchmark.
For enterprise buyers, the missing pricing is the showstopper. You cannot budget a generation tool with no published rate card, no API tier, and no service-level commitment. VentureBeat put it bluntly: "Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption."
FLUX 3 is the first credible answer to the question the frontier labs have been ducking: can one model actually learn image, video, audio, and action together well enough to ship? BFL says yes, and the early visual evidence agrees. The architecture argument — that joint training beats stitched pipelines — just got its first real demonstration outside a research paper.
It is also the first 2026 release where the open-weight story is the suspense, not the headline. The FLUX 3 Dev release is the most important open-source drop of the back half of 2026.
For everyone else: apply for Early Access if you have a video workload that needs native audio. The product is real. The pricing is not. The architecture is the story.
— Mr. Technology