FLUX 3 — Multimodal AI for Video, Image & Audio

Black Forest Labs' next-generation foundation model. FLUX 3 learns images, video, and audio together — so motion, sound, and structure stay more coherent. Generate clips with native audio (commonly up to about 20 seconds), continue from images or video, and prepare for stronger image synthesis with complex prompts and multilingual typography.

Available on SeedDance — FLUX 3 Video (early access)

FLUX 3 multimodal overview

One Backbone for Visual Intelligence

FLUX 3 is Black Forest Labs' multimodal foundation model announced in 2026. Unlike earlier FLUX releases that focused mainly on images, FLUX 3 is trained jointly on images, video, and audio in a unified Self-Flow architecture — aiming for a stronger shared understanding of how objects hold together, how they move, and how events sound. Product surfaces roll out as FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and later FLUX 3 Dev (open weights).

Video + Native Audio

Generate diverse clips with synchronized speech, effects, and ambience in one pass. Official early materials highlight single-generation lengths commonly up to about 20 seconds (resolution and mode depend on the provider API).

Text, Image & Video Inputs

Start from a prompt (text-to-video), pin images as frames or references (image-to-video), or continue an existing clip (video continuation) while carrying characters and context forward.

Stronger Image Synthesis

FLUX 3 Image targets better complex-prompt following and high-accuracy multilingual text in-frame — across flexible styles, aspect ratios, and resolutions — with early access rolling out after video.

World-Consistent Generation

Joint training across modalities is designed so sound matches impact, motion respects structure, and multi-shot sequences stay more coherent for longer storytelling drafts.

Built for Coherent Motion, Sound & Style

FLUX 3 is positioned for creators and teams who need more than a single beautiful frame — short narrative video with audio, reference-driven continuity, and image work that follows dense prompts.

Native audio generation pairs lipsync-oriented dialogue with effects and ambience tied to physical events — reducing the need to assemble a separate soundtrack for many drafts.

FLUX 3 native audio

FLUX 3 Capability Overview

Core capabilities of FLUX 3 Video on SeedDance. Exact options match the AI video generator (duration, resolution, draft, audio, and inputs).

Text-to-Video

Describe scene, camera, style, and dialogue in natural language to generate a short video with optional native audio.

Image-to-Video

One image opens the clip pixel-for-pixel; two set start and end; three to ten act as a storyboard (explicit duration required). Images and start video cannot be combined.

Video Continuation

Upload a short mp4 (≤15s, ≤50MB) as start_video to continue from its final frames. Cannot be combined with image inputs.

Draft Preview

Generate a fast, lower-cost 720p draft to iterate on prompts before a full-quality 720p or 1080p run.

Synchronized Audio

Audio is on by default (ambience, speech, effects). Turn it off for a silent clip.

Image Synthesis & Editing

FLUX 3 Image aims for wide style range, flexible framing, and better in-image text — with API early access rolling out after the video tier.

Frequently Asked Questions

Common questions about FLUX 3 and using related models on SeedDance.








Create with FLUX 3 Video on SeedDance

Open the AI video generator, select FLUX 3, and generate clips with optional native audio — from text, images, or a start video.