Text-to-video asks a model to invent a composition, a colour palette, a subject and a motion, all at once, from a sentence. Then you pay for the result whether or not it resembles what you pictured.
Image-to-video removes three of those four variables. You fix the frame first — cheaply, in an image model where iteration costs a fraction of a video generation — and then ask only for movement.
For most production work this is simply the better workflow, and it is meaningfully cheaper.
Why it works better
Composition stops being random. You have already decided where the subject sits, what is in the background and how the frame is cropped. The model is no longer guessing.
Iteration gets cheap. Getting a composition right can take a dozen attempts. Doing that in an image model rather than a video model is a large cost difference — you are comparing cents against dollars per attempt.
Consistency becomes possible. Animating the same source image several times with different motion prompts gives you shots that genuinely belong together. Generating them independently from text does not.
You can use assets you already own. Product photography, a client's brand imagery, a frame from existing footage, a still you shot yourself. That is often the only workable route for commercial work.
Choosing a frame that will animate well
Not every image animates well, and the failures are predictable.
Leave room for the motion. A subject filling the frame edge to edge has nowhere to move without leaving it. If you want a walk, leave space in the direction of travel.
Give it depth. A flat, single-plane image produces flat, unconvincing motion. Foreground, midground and background give the model parallax to work with, which is most of what makes a camera move read as real.
Keep the subject unambiguous. If the model cannot tell where the person ends and the background begins, it will animate the boundary badly. Reasonable subject/background separation matters more than resolution.
Watch out for hands, text and faces at small sizes. These are where artefacts show first. If a hand is a few dozen pixels across, expect it to misbehave; crop tighter or accept it.
Start sharp. Motion blur and low resolution in the source compound during animation. Feed it the best version you have.
Prompt the motion, not the picture
This is the mistake that ruins most image-to-video attempts.
The model can already see your image. Re-describing it invites the model to re-interpret it — and the composition you carefully built gets rebuilt into something else.
Do not write:
A woman in a red coat standing on a bridge over a river at sunset, city skyline behind her, cinematic lighting
The model has all of that. You are asking it to reconsider.
Write:
Slow push in. Her coat and hair move in a light breeze; water flows beneath the bridge. Subtle camera drift, everything else still.
Two sentences, entirely about movement.
A vocabulary for motion
Camera moves
- push in / pull out — the workhorse. Small amounts read as expensive.
- pan left / pan right — rotation from a fixed position.
- orbit / arc around the subject — excellent for products, demanding for the model.
- tilt up / tilt down — reveals.
- handheld, subtle drift — adds life without committing to a direction.
- locked off, static camera — say it explicitly if you want no camera motion. Otherwise you will usually get some.
Subject motion
- hair and clothing move in the wind
- steam rises, water flows, leaves drift
- she turns her head slowly toward camera
- he blinks and takes a breath
Intensity
- subtle, slow, gentle, barely perceptible — reach for these first.
- dramatic, fast, sweeping — usually the thing that broke your generation.
The reliable pattern
[Camera move]. [One subject motion]. [What stays still]. [Intensity].
For example:
Slow orbit to the right around the watch. Light moves across the polished case as the camera travels. The watch itself stays perfectly still on the stand. Smooth, slow, steady.
That last sentence — naming what should not move — is disproportionately effective. Models have a bias towards adding motion everywhere, and explicitly pinning things down counteracts it.
Which model to use
Wan 2.6 and PixVerse V6 are image-to-video specialists. Both tend to hold the source composition faithfully, which is the whole point of this workflow.
PixVerse V6 will generate as short as 1 second, which is unusual and useful for cutaways and inserts.
Seedance 2.5, Seedance 2.0, Veo 3.1, Kling 3.0, Gemini Omni and Runway Gen-4 all accept a starting image while also doing text-to-video. Several of them additionally support a last frame, so you can specify where the motion should end as well as where it starts — that is the tool for a controlled reveal or a match cut.
If you need a character to stay recognisable across several shots, use a model with reference-image support: Seedance 2.5, Seedance 2.0, Veo 3.1 or Gemini Omni.
A working pipeline
1. Generate or choose the source image. If generating, iterate here until the frame is genuinely right. This is the cheap part — do not rush it.
2. Upload it and pick a model. Image-to-video specialist for fidelity; general model if you also want generated audio.
3. Prompt only the motion. Camera move, one subject motion, what stays still, intensity.
4. Draft short and low. Judge at the shortest duration and lowest resolution. Motion problems are perfectly visible at 480p.
5. Diagnose honestly. If the composition drifted, your prompt is describing the scene — cut it back to motion only. If the motion is mushy, ask for less of it. If the subject deforms, the source may not have enough separation between subject and background.
6. Final pass. Once the motion is right, run it once at full resolution and duration.
Common failures and their fixes
| What you see | Why | Fix |
|---|---|---|
| Composition changes entirely | Prompt re-describes the scene | Prompt motion only |
| Subject melts or deforms | Motion requested is too large | Ask for subtler movement |
| Nothing moves | Prompt too vague | Name a specific thing that moves |
| Everything moves, including things that shouldn't | No anchor given | Explicitly state what stays still |
| Background warps during a camera move | Not enough depth in the source | Choose an image with clearer foreground/background separation |
| Face looks wrong | Face too small in frame | Crop tighter, or accept it at that scale |
Why this is also the cheaper route
Composition iteration is the expensive part of AI video, and image-to-video moves that iteration into a medium that costs a fraction as much per attempt. A workflow of ten image generations plus two video generations costs dramatically less than twelve video generations — and tends to produce a better result, because you were solving one problem at a time.
Upload a still and animate it on TurboMax AI — image-to-video is available on Wan 2.6, PixVerse V6, Seedance, Veo 3.1, Kling 3.0 and more, all from one prepaid balance with no subscription.