ByteDance Seed's full-scene audio model on SeedDance. Orchestrate multi-character dialogue, emotional delivery, ambience, and sound effects from a single prompt — up to about two minutes of audio per generation.
Available on SeedDance platform

Seed Audio 1.0 is a multimodal audio creation model from the ByteDance Seed team. Unlike conventional text-to-speech that outputs a single voice track, it generates complete sound scenes — dialogue plus the acoustic world around it. On SeedDance you can create from a text prompt (up to 1,500 characters), optionally conditioned with up to three reference audio clips or one reference image, for outputs up to about two minutes.
Traditional TTS turns text into one voice. Seed Audio 1.0 targets the full soundscape: dialogue, ambience, effects, and scene texture layered together. Describe a scene in natural language and get a listenable mix instead of stitching multiple tools.
Upload up to three reference clips and tag them in the prompt with @Audio1, @Audio2, and @Audio3 to guide voice style. Or use one reference image for mood when you are not using audio references — image and audio references cannot be combined in the same generation.
Generate conversations with distinct speakers, each with its own timbre and emotional arc. Useful for scripted podcasts, training scenarios, and character-driven storytelling without recording multiple voice actors.
Environmental ambience and action-matched sound effects can be generated alongside speech in a shared scene context — reducing separate SFX hunts and rough mixes for prototypes and short-form content.
Capabilities available in the SeedDance AI Audio Generator — with clear limits for prompts, references, and output length.

Capabilities exposed in the SeedDance AI Audio Generator for Seed Audio 1.0.
Describe characters, setting, mood, and pacing in natural language (up to 1,500 characters). The model renders a complete audio scene rather than a flat narration track.
Upload up to three reference clips (typically up to 30 seconds and 10 MB each) and reference them with @Audio1, @Audio2, @Audio3 for voice style and character casting.
Supply a single reference image (JPEG, PNG, or WebP) to influence mood when audio references are not used. Image and audio references are mutually exclusive.
Assign distinct voices to multiple speakers within one generation for scripted conversations, interviews, and narrative exchanges.
Generate ambient sound design with dialogue — rain, footsteps, city noise, mechanical hum, and other scene layers described in the prompt.
Specify when lines should enter in your prompt. Official timing control for character dialogue is available at about 100 ms intervals — useful for dubbing and timed scripts.
Generate expressive scene audio across many languages (officially 20+), including Chinese, English, Japanese, Korean, Spanish, and more — useful for localized drafts.
Produce up to about two minutes of audio in a single SeedDance generation — enough for podcast intros, ad spots, and short dramatic scenes.
Everything you need to know about Seed Audio 1.0 on SeedDance.
Try full-scene AI audio — multi-character dialogue, ambience, and effects from text and optional references, up to about two minutes per generation.