AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

FLUX 3

Black Forest Labs·

Videomultimodalcloud

Overview

Unified multimodal flow model that generates image, video and audio from a single backbone and can be extended to predict robot actions. Rollout is staged: FLUX 3 Video with native synchronised audio (clips up to 20 seconds, aspect ratios from 9:16 to 21:9) and the action model are in application-based early access, image generation is announced for the following weeks, and an open-weight FLUX 3 Dev backbone for later in the year. No public API tier, pricing, parameter count or benchmark methodology published at launch. FLUX 3 Video now ranks #2 in the Text-to-Video Arena at ~1,496 Elo, behind Gemini Omni Flash (~1,512) and ahead of Dreamina Seedance 2.0 (~1,478) — the first independent quality signal for the model.

Capabilities and innovations

Text-to-video (up to 20s)Native synchronised audioImage-to-videoVideo-to-video from reference clipKeyframe-to-video transitionsRobot action predictionSingle flow backbone for image, video, audio and actionVideo and native audio generated jointlyText-to-Video Arena #2 at intake (~1,496 Elo)

Benchmarks

No tracked academic benchmark applies to video models; the timeline carries no scores and no leadership tracking for them.

  • GPQA Diamond: not reported by the vendor
  • Humanity’s Last Exam (no tools): not reported by the vendor
  • SWE-bench Verified: not reported by the vendor
  • SWE-bench Pro: not reported by the vendor
  • MMLU: not reported by the vendor
  • MMLU-Pro: not reported by the vendor
  • Terminal-Bench 2.x: not reported by the vendor
  • MMMU-Pro: not reported by the vendor

Links

More from Black Forest Labs

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.