Seed Audio 1.0 — Full-Scene AI Audio in One GenerationFull-Scene Audio

ByteDance Seed's full-scene audio model on SeedDance. Orchestrate multi-character dialogue, emotional delivery, ambience, and sound effects from a single prompt — up to about two minutes of audio per generation.

Available on SeedDance platform

Seed Audio 1.0 Overview

What is Seed Audio 1.0

Seed Audio 1.0 is a multimodal audio creation model from the ByteDance Seed team. Unlike conventional text-to-speech that outputs a single voice track, it generates complete sound scenes — dialogue plus the acoustic world around it. On SeedDance you can create from a text prompt (up to 1,500 characters), optionally conditioned with up to three reference audio clips or one reference image, for outputs up to about two minutes.

Beyond Text-to-Speech

Traditional TTS turns text into one voice. Seed Audio 1.0 targets the full soundscape: dialogue, ambience, effects, and scene texture layered together. Describe a scene in natural language and get a listenable mix instead of stitching multiple tools.

Reference Audio or Image

Upload up to three reference clips and tag them in the prompt with @Audio1, @Audio2, and @Audio3 to guide voice style. Or use one reference image for mood when you are not using audio references — image and audio references cannot be combined in the same generation.

Multi-Role Dialogue & Emotion

Generate conversations with distinct speakers, each with its own timbre and emotional arc. Useful for scripted podcasts, training scenarios, and character-driven storytelling without recording multiple voice actors.

Ambience & SFX in One Pass

Environmental ambience and action-matched sound effects can be generated alongside speech in a shared scene context — reducing separate SFX hunts and rough mixes for prototypes and short-form content.

What You Can Do with Seed Audio 1.0 on SeedDance

Capabilities available in the SeedDance AI Audio Generator — with clear limits for prompts, references, and output length.

Coordinate dialogue, ambience, and sound cues under one scene description. A late-night convenience-store radio drama can include whispered lines, fluorescent hum, door chimes, and tense underscore from a single instruction — faster drafts without a full studio stack.

One Prompt One Mix

SeedDance Feature Set for Seed Audio 1.0

Capabilities exposed in the SeedDance AI Audio Generator for Seed Audio 1.0.

Text-to-Audio Scene Generation

Describe characters, setting, mood, and pacing in natural language (up to 1,500 characters). The model renders a complete audio scene rather than a flat narration track.

Reference Audio Conditioning

Upload up to three reference clips (typically up to 30 seconds and 10 MB each) and reference them with @Audio1, @Audio2, @Audio3 for voice style and character casting.

Optional Image Reference

Supply a single reference image (JPEG, PNG, or WebP) to influence mood when audio references are not used. Image and audio references are mutually exclusive.

Multi-Character Dialogue

Assign distinct voices to multiple speakers within one generation for scripted conversations, interviews, and narrative exchanges.

Ambience & Environmental FX

Generate ambient sound design with dialogue — rain, footsteps, city noise, mechanical hum, and other scene layers described in the prompt.

Prompt-Level Dialogue Timing

Specify when lines should enter in your prompt. Official timing control for character dialogue is available at about 100 ms intervals — useful for dubbing and timed scripts.

Multilingual Scene Audio

Generate expressive scene audio across many languages (officially 20+), including Chinese, English, Japanese, Korean, Spanish, and more — useful for localized drafts.

Output up to ~2 Minutes

Produce up to about two minutes of audio in a single SeedDance generation — enough for podcast intros, ad spots, and short dramatic scenes.

Frequently Asked Questions

Everything you need to know about Seed Audio 1.0 on SeedDance.









Start Creating with Seed Audio 1.0 on SeedDance

Try full-scene AI audio — multi-character dialogue, ambience, and effects from text and optional references, up to about two minutes per generation.