Dialogue with lip-sync
“A street food chef looks up at the camera mid-service and says “best noodles in the city, trust me” with a grin, woks flaming behind, handheld framing.”
video model
Happy Horse 1.1 is Alibaba’s flagship video model — a long-running #1 on the Artificial Analysis Video Arena — generating 3 to 15 second clips at 1080p with native audio and multilingual lip-sync.
$0.18 / second at provider cost for members — see the full price breakdown
What you need to start
Describe the scene, motion, and any dialogue — or add a start image to animate.
Original results
These examples use media produced for Made. Prompts are included when the production record allows them.
“A street food chef looks up at the camera mid-service and says “best noodles in the city, trust me” with a grin, woks flaming behind, handheld framing.”
“Animate the scene: steam rises, the chef tosses the wok once, ambient market sound builds, the camera holds steady.”
“Horses gallop through shallow water at sunrise, spray catching the light in slow motion, hooves drumming, a low tracking shot alongside.”

“Use this image as the opening composition.”

“Set the scene as a still before describing its motion and sound.”

“Plan the shot as a still, then hand the motion to the model.”
These are the controls exposed by Made, which may be narrower than a provider’s full API.
Describe the subject, motion, and camera; put spoken lines in quotes for lip-synced delivery.
Upload a still to animate at 1080p while keeping its composition.
Pick one of nine aspect ratios and a duration between 3 and 15 seconds.
Check lip-sync, audio, and motion continuity, then refine one instruction at a time.
video
FLUX 3 is Black Forest Labs’ first video model, generating 5 to 20 second clips with natively synchronized audio from a text prompt or a start image, in seven aspect ratios up to 2:1.
Explore model →video
Kling 3 Pro is the benchmark tier of Kling AI’s third-generation video model, generating 3 to 15 second clips with audio on by default and the cinematic motion the Kling line is known for.
Explore model →video
WAN 3.0 is the standard tier of Alibaba’s WAN 3.0 video line, generating 2 to 30 second clips at 1080p with audio included — half the per-second rate of WAN 3.0 Prime.
Explore model →Turn any image into a captivating video. Upload a photo, describe the motion you want, and watch AI bring it to life with realistic movement, cinematic camera work, and professional quality.
Open tool →Create scroll-stopping product photos and marketing visuals in seconds. Describe your shot, pick your style, and get professional images ready for your store, social feed, or ad campaign.
Open tool →It is Alibaba’s flagship video model, which took the #1 Elo spot on the Artificial Analysis Video Arena in 2026, generating 1080p clips with native audio and multilingual lip-sync.