HappyHorse 1.1 — Native Audio, Stronger Motion, Up to 9 References

Alibaba ATH's production upgrade to HappyHorse. Create 3–15 second clips at 720p or 1080p with joint audio-video generation, multilingual lip sync, and reference-to-video consistency across up to nine images — live on SeedDance.

Powered by Alibaba ATH · Available on SeedDance

HappyHorse 1.1 AI video overview on SeedDance

What is HappyHorse 1.1

HappyHorse 1.1 is Alibaba ATH's (Taotian Group) upgraded AI video generation model, released in June 2026. Built on a unified ~15B-parameter Transformer that generates video and synchronized audio in one pass, it improves five production dimensions over HappyHorse 1.0: motion expressiveness, subject consistency, instruction following, visual quality, and audio-visual sync — while keeping the same 3–15 second / 720p–1080p specs for easy migration.

Joint Audio-Video Generation

Dialogue, ambient sound, Foley, and music cues are generated with the picture in a single pass — so lip sync, pacing, and on-screen action stay aligned without a separate dubbing pipeline.

Up to 9 Reference Images (R2V)

Reference-to-video locks characters, products, packaging, and environments across clips. Upload 1–9 images and bind them in the prompt for stronger brand and IP consistency than text-only generation.

Stronger Motion & Prompt Following

Version 1.1 targets smoother action (dance, sports, product spins) and sharper adherence to camera, lighting, and shot-by-shot instructions — fewer retries for production teams.

Multilingual Lip Sync

Phoneme-aware lip sync across major languages including English, Mandarin, Cantonese, Japanese, Korean, German, and French — useful for localized ads and talking-head variants.

Why Choose HappyHorse 1.1 on SeedDance

Built for short-form production workflows — social ads, e-commerce demos, short drama beats, and multilingual marketing — with clear specs and SeedDance credits.

Same duration and resolution envelope as 1.0, with clearer gains in motion, multi-reference consistency, instruction following, visual detail, and audio-visual sync. New projects should start on 1.1; migrating existing 1.0 prompts is usually straightforward.

HappyHorse 1.1 upgrades

HappyHorse 1.1 Feature Set on SeedDance

Capabilities available in the SeedDance AI video generator for HappyHorse 1.1.

Text-to-Video (T2V)

Describe scenes, camera moves, performance, and dialogue in natural language. Supports detailed shot-by-shot direction for multi-beat short narratives.

Image-to-Video (I2V)

Animate a still while preserving subject and style. Add an optional motion prompt to guide camera and action.

Reference-to-Video (R2V)

Upload 1–9 reference images to lock characters, products, and environments. Reference subjects in the prompt for multi-subject consistency.

Native Synchronized Audio

Joint generation of dialogue, ambience, and effects with the video — wrap spoken lines in quotes to improve lip sync alignment.

720p & 1080p Output

Choose 720p for faster, lower-cost drafts or 1080p for sharper delivery. Credits scale by duration and quality on SeedDance.

3–15 Second Duration

Any integer length from 3 to 15 seconds (default 5). Linear per-second billing — longer clips cost proportionally more.

Flexible Aspect Ratios

Support for common social and cinematic frames including 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, and 21:9.

Production-Friendly Value

On SeedDance, HappyHorse 1.1 starts from 50 credits for a 5-second 720p generation (100 credits at 1080p), with image-to-video and reference-to-video on the same credit curve.

Frequently Asked Questions

Everything you need to know about HappyHorse 1.1 on SeedDance.










Start Creating with HappyHorse 1.1 Today

Try Alibaba's HappyHorse 1.1 on SeedDance — joint audio-video generation, up to 9 reference images, 720p/1080p, and 3–15 second clips for ads, short drama, and product storytelling.