FLUX 3
Black Forest Labs·
Overview
Unified multimodal flow model that generates image, video and audio from a single backbone and can be extended to predict robot actions. Rollout is staged: FLUX 3 Video with native synchronised audio (clips up to 20 seconds, aspect ratios from 9:16 to 21:9) and the action model are in application-based early access, image generation is announced for the following weeks, and an open-weight FLUX 3 Dev backbone for later in the year. No public API tier, pricing, parameter count or benchmark methodology published at launch. FLUX 3 Video now ranks #2 in the Text-to-Video Arena at ~1,496 Elo, behind Gemini Omni Flash (~1,512) and ahead of Dreamina Seedance 2.0 (~1,478) — the first independent quality signal for the model.
Capabilities and innovations
Benchmarks
No tracked academic benchmark applies to video models; the timeline carries no scores and no leadership tracking for them.
- GPQA Diamond: not reported by the vendor
- Humanity’s Last Exam (no tools): not reported by the vendor
- SWE-bench Verified: not reported by the vendor
- SWE-bench Pro: not reported by the vendor
- MMLU: not reported by the vendor
- MMLU-Pro: not reported by the vendor
- Terminal-Bench 2.x: not reported by the vendor
- MMMU-Pro: not reported by the vendor
Links
More from Black Forest Labs
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.