返回课程
Lesson 04Intermediate9 分钟

Image-to-video: animating stills with control

How to use HappyHorse-1.0 image-to-video: choosing source images that animate well, writing motion-only prompts, and the reference-to-video mode for character consistency.

更新于 2026-07-17

Text-to-video starts from nothing; image-to-video (i2v) starts from a frame you already trust. That makes i2v the control freak's mode: composition, character design, color palette — all locked before generation begins. HappyHorse-1.0 ranked #1 in image-to-video at its arena debut, and it remains the mode where the model feels most reliable.

When to reach for i2v

  • Brand and product work — the product must look exactly like the product. Generate or photograph the hero frame first, then animate it.
  • Character continuity — the same face across multiple clips. Animate the same portrait with different motion prompts.
  • Art direction — you have a style frame from a designer (or an image model) and want it moving without re-rolling the look.
  • Rescuing t2v near-misses — a text-to-video clip with a perfect first frame but wrong motion? Extract the frame, switch to i2v, re-prompt the motion.

Choosing source images that animate well

The model animates what's plausible from the pixels it's given. Images that work:

  • Clear subject with implied motion — a runner mid-stride, steam rising, fabric caught in wind. The image suggests its own next frame.
  • Room to move — headroom above a jumping subject, road ahead of a car. If the subject fills 100% of the frame, motion has nowhere to go.
  • Coherent lighting — the model extends your lighting; contradictory light sources produce shimmer.

Images that fight you: busy collages, heavy text overlays, extreme close-ups, and stylized images with inconsistent perspective. Text in the source tends to warp once things start moving — keep titles for post.

Writing the motion prompt

In i2v mode your prompt should describe what changes, not what's already visible. The image is the noun; the prompt is the verb.

Weak (re-describes the image):

A woman in a red coat stands on a bridge in the rain

Strong (directs the motion):

She turns toward the camera and smiles as the wind lifts her hair, umbrella tilting; slow push-in, rain intensifies, distant traffic hum

Structure it as: subject motion + secondary motion (hair, weather, background) + camera movement + audio cue. Everything visual-style-related is already decided by your image.

Reference-to-video (r2v)

HappyHorse exposes a third mode, r2v, where the input image acts as a reference rather than a literal first frame — the model preserves the subject's identity while re-staging it in a new scene described by your prompt. Use i2v when the composition is sacred; use r2v when the character is sacred but the scene changes ("same mascot, now surfing"). r2v availability varies by host platform; the official API exposes it directly (Lesson 7).

A practical pipeline

  1. Generate 4–8 candidate stills in an image model (or shoot a photo).
  2. Pick the frame with the best composition and animation headroom.
  3. Animate at 720p / 5s with a motion-only prompt.
  4. Iterate on motion — the image never changes, so every attempt is a clean A/B on your prompt.
  5. Upscale the winner: re-render at 1080p, or extend duration if the motion sustains.

This pipeline is cheaper than pure t2v iteration because you stop paying for composition re-rolls — you're only iterating the part that changes.

Audio still comes free

Even in i2v, HappyHorse generates synchronized audio from scene context: animate a beach still and you get surf; a concert still gets crowd noise. Mention the soundscape in the prompt when it matters ("waves crashing, gulls overhead") — silent-looking sources sometimes produce timid audio otherwise.

Next: the mode where that audio engine becomes the whole point — dialogue and lip-sync.