FLUX 3 — Multimodal AI for Video, Image & Audio

Black Forest Labs' next-generation foundation model. FLUX 3 learns images, video, and audio together — so motion, sound, and structure stay more coherent. Generate clips with native audio (commonly up to about 20 seconds), and continue from images or video when the video tier is live on SeeVid.

Coming Soon on SeeVid — FLUX 3 Video

FLUX 3 video generation overview

Multimodal Video from Black Forest Labs

FLUX 3 is Black Forest Labs' multimodal foundation model announced in 2026. Unlike earlier FLUX releases that focused mainly on images, FLUX 3 is trained jointly on images, video, and audio in a unified Self-Flow architecture — aiming for a stronger shared understanding of how objects hold together, how they move, and how events sound. On SeeVid, the planned product surface is FLUX 3 Video (text, image, and video inputs with optional native audio).

Video + Native Audio

Generate diverse clips with synchronized speech, effects, and ambience in one pass. Official early materials highlight single-generation lengths commonly up to about 20 seconds (resolution and mode depend on the provider API).

Text, Image & Video Inputs

Start from a prompt (text-to-video), pin images as frames or references (image-to-video), or continue an existing clip (video continuation) while carrying characters and context forward.

720p & 1080p Output

Planned SeeVid options include 720p and 1080p, plus a fast draft preview for iterating on prompts before a full-quality run.

World-Consistent Generation

Joint training across modalities is designed so sound matches impact, motion respects structure, and multi-shot sequences stay more coherent for longer storytelling drafts.

Built for Coherent Motion, Sound & Style

FLUX 3 is positioned for creators and teams who need more than a single beautiful frame — short narrative video with audio and reference-driven continuity.

Native audio generation pairs lipsync-oriented dialogue with effects and ambience tied to physical events — reducing the need to assemble a separate soundtrack for many drafts.

FLUX 3 native audio

FLUX 3 Video Capability Overview

Planned capabilities for FLUX 3 Video on SeeVid. Exact options will match the AI video generator when the model is restored (duration, resolution, draft, audio, and inputs).

Text-to-Video

Describe scene, camera, style, and dialogue in natural language to generate a short video with optional native audio.

Image-to-Video

One image opens the clip pixel-for-pixel; two set start and end; three to ten act as a storyboard (explicit duration required). Images and start video cannot be combined.

Video Continuation

Upload a short mp4 (≤15s, ≤50MB) as start_video to continue from its final frames. Cannot be combined with image inputs.

Draft Preview

Generate a fast, lower-cost 720p draft to iterate on prompts before a full-quality 720p or 1080p run.

Synchronized Audio

Audio is on by default (ambience, speech, effects). Turn it off for a silent clip.

Duration 5–20s or Auto

Pick a whole number of seconds from 5 to 20, or auto so the model fits the content. Storyboards with three or more images need an explicit duration.

Frequently Asked Questions

Common questions about FLUX 3 and SeeVid availability.








FLUX 3 Video is Coming Soon on SeeVid

We are preparing FLUX 3 Video for text, image, and start-video workflows with optional native audio. Check back when early access returns.