Image-to-video: animating stills with control
How to use HappyHorse-1.0 image-to-video: choosing source images that animate well, writing motion-only prompts, and the reference-to-video mode for character consistency.
更新于 2026-07-17
Text-to-video starts from nothing; image-to-video (i2v) starts from a frame you already trust. That makes i2v the control freak's mode: composition, character design, color palette — all locked before generation begins. HappyHorse-1.0 ranked #1 in image-to-video at its arena debut, and it remains the mode where the model feels most reliable.
When to reach for i2v
- Brand and product work — the product must look exactly like the product. Generate or photograph the hero frame first, then animate it.
- Character continuity — the same face across multiple clips. Animate the same portrait with different motion prompts.
- Art direction — you have a style frame from a designer (or an image model) and want it moving without re-rolling the look.
- Rescuing t2v near-misses — a text-to-video clip with a perfect first frame but wrong motion? Extract the frame, switch to i2v, re-prompt the motion.
Choosing source images that animate well
The model animates what's plausible from the pixels it's given. Images that work:
- Clear subject with implied motion — a runner mid-stride, steam rising, fabric caught in wind. The image suggests its own next frame.
- Room to move — headroom above a jumping subject, road ahead of a car. If the subject fills 100% of the frame, motion has nowhere to go.
- Coherent lighting — the model extends your lighting; contradictory light sources produce shimmer.
Images that fight you: busy collages, heavy text overlays, extreme close-ups, and stylized images with inconsistent perspective. Text in the source tends to warp once things start moving — keep titles for post.
Writing the motion prompt
In i2v mode your prompt should describe what changes, not what's already visible. The image is the noun; the prompt is the verb.
Weak (re-describes the image):
A woman in a red coat stands on a bridge in the rain
Strong (directs the motion):
She turns toward the camera and smiles as the wind lifts her hair, umbrella tilting; slow push-in, rain intensifies, distant traffic hum
Structure it as: subject motion + secondary motion (hair, weather, background) + camera movement + audio cue. Everything visual-style-related is already decided by your image.
Reference-to-video (r2v)
HappyHorse exposes a third mode, r2v, where the input image acts as a reference rather than a literal first frame — the model preserves the subject's identity while re-staging it in a new scene described by your prompt. Use i2v when the composition is sacred; use r2v when the character is sacred but the scene changes ("same mascot, now surfing"). r2v availability varies by host platform; the official API exposes it directly (Lesson 7).
A practical pipeline
- Generate 4–8 candidate stills in an image model (or shoot a photo).
- Pick the frame with the best composition and animation headroom.
- Animate at 720p / 5s with a motion-only prompt.
- Iterate on motion — the image never changes, so every attempt is a clean A/B on your prompt.
- Upscale the winner: re-render at 1080p, or extend duration if the motion sustains.
This pipeline is cheaper than pure t2v iteration because you stop paying for composition re-rolls — you're only iterating the part that changes.
Audio still comes free
Even in i2v, HappyHorse generates synchronized audio from scene context: animate a beach still and you get surf; a concert still gets crowd noise. Mention the soundscape in the prompt when it matters ("waves crashing, gulls overhead") — silent-looking sources sometimes produce timid audio otherwise.
Next: the mode where that audio engine becomes the whole point — dialogue and lip-sync.