Most advice on AI video prompting is a list of magic words. Add "cinematic", add "8k", add "masterpiece, highly detailed, award winning".
That advice is a holdover from early image models, where those tokens genuinely nudged output towards a particular training subset. Modern video models are far better at instruction-following, and the magic words now mostly consume prompt budget without changing the result.
What works instead is describing a shot the way you would describe it to a camera operator.
The six components
Almost every prompt that works contains some of these, in roughly this order of importance.
1. Subject — who or what is in frame, described concretely. "A woman in her sixties with grey hair tied back, wearing a wax jacket" beats "a person".
2. Action — one thing happening. "pours coffee from a steel flask". Not "pours coffee, looks up, smiles, walks away". More on this below, because it is the single most common failure.
3. Camera — where the camera is and what it does. "handheld medium shot, slow push in". This is the component people most often omit, and it is the one that most reliably improves output. Without it the model picks, and it usually picks a static wide.
4. Setting — where, and what time. "on a fishing boat at dawn, harbour lights still on".
5. Lighting — the single highest-leverage aesthetic control. "low golden sunlight from camera left, deep shadows". Lighting does more for whether footage reads as "cinematic" than any quality adjective will.
6. Style — only if you want something specific. "shot on 16mm, visible grain, muted colours". Skip it if you want photoreal; the model defaults there anyway.
Put together:
Handheld medium shot, slow push in. A woman in her sixties with grey hair tied back, wearing a wax jacket, pours coffee from a steel flask on a fishing boat at dawn. Low golden sunlight from camera left, deep shadows, harbour lights still on in the background. Shot on 16mm, visible grain, muted colours.
Fifty-two words. Every one of them describes something a camera could record.
The mistake nearly everyone makes
More than one action in a prompt.
Video models generate a continuous motion field over a few seconds. Ask for a sequence of events and you get all of them smeared together, in the wrong order, or only the first one.
This does not work:
A man walks into the kitchen, opens the fridge, takes out a beer, closes the door and walks out.
That is five shots. In a 5-second generation you will get a man doing something vaguely kitchen-shaped with limbs in the wrong places.
This works:
Medium shot, static camera. A man opens a fridge door and reaches inside, warm interior light spilling onto his face in a dark kitchen.
One action. One moment. If you need the full sequence, generate it as separate shots and cut them together — which is what a film crew would do anyway.
Why "cinematic 8k masterpiece" does nothing
Those words are not filmable. "Cinematic" describes a feeling that comes out of specific choices: shallow depth of field, motivated lighting, deliberate camera movement, a particular colour treatment.
If you want the feeling, specify the choices:
| Instead of | Write |
|---|---|
| cinematic | shallow depth of field, anamorphic lens flare, slow dolly in |
| 8k, highly detailed | close-up, sharp focus on the subject's hands |
| dramatic | single hard key light from below, deep shadows, high contrast |
| beautiful lighting | golden hour backlight, lens haze, warm rim light on hair |
| epic | wide establishing shot, low camera angle, subject small against the landscape |
The right-hand column is longer. It is also the difference between a model guessing and a model executing.
Debugging a prompt that will not work
When output is consistently wrong, resist the urge to add more words. Adding usually makes it worse. Instead:
Cut back to the subject and one action. Generate. If that is right, the problem is in what you added — put it back one element at a time. If it is still wrong, the model may not be able to do this. Try a different one.
Check whether you are describing something abstract. "A sense of loss", "the feeling of nostalgia", "representing freedom" — none of these are photographable. Translate to the image you actually have in your head: an empty chair, a dusty window, a bird leaving a wire.
Move the important thing earlier. Long prompts lose their tail. If the model keeps ignoring an element, move it to the front.
Check the negatives are earning their place. Only useful on models that support them, and only for artefacts you have actually observed. Pre-emptively listing "no extra fingers, no distortion, no watermark" is at best neutral and at worst counterproductive.
Try image-to-video instead. If the composition keeps coming out wrong, stop fighting it in text. Generate a still until the frame is right — cheaper per attempt — then animate that. You have removed composition from the list of things the video model has to get right.
Drafting cheaply
The iteration loop costs money, so make each loop as cheap as possible:
- Iterate at the lowest resolution the model offers. You are judging composition and motion, and both are legible at 480p.
- Use the shortest duration while testing. If the motion is right at 4 seconds it will be right at 10.
- Use a fast model variant — Veo 3.1 Fast rather than Veo 3.1 — while the prompt is still moving.
- Turn audio off while iterating.
- Only when the prompt is locked, run it once at full resolution, full duration, audio on.
A prompt that takes eight attempts costs roughly a quarter as much drafted at 480p as it does at 1080p. That is a bigger saving than any pricing plan will give you.
Prompts worth stealing
Product, clean and commercial
Slow orbit around a matte black wireless speaker on a concrete surface. Soft top light, single hard rim light from behind. Shallow depth of field, seamless light grey background. Macro detail on the speaker grille.
Character, cinematic
Close-up, slow push in. A young man in a rain-soaked denim jacket looks up as neon signs reflect in the puddles around him. Night, heavy rain, cyan and magenta practical lights. Shallow focus, anamorphic flare, 35mm.
Landscape, establishing
Static wide shot, locked off. Mist moving through a pine valley at first light, layered ridgelines fading into haze. Cool blue shadows, warm sun catching the highest trees. No camera movement.
Food
Overhead static shot. Steam rising from a bowl of ramen on a dark wooden table, chopsticks lifting noodles into frame. Warm side light from camera right, deep shadows, shallow depth of field.
Note what they have in common: camera first, one action, specific light, no adjective doing work a noun could do better.
The short version
Write the shot, not the vibe. One action per generation. Always specify the camera. Lighting does the heavy lifting. Draft cheap, finish expensive. And when a prompt will not work, cut it back rather than piling more on.
Every model on TurboMax AI shows its credit cost before you generate, and failed generations are refunded automatically — so iterating on a prompt costs you exactly the attempts that produced something.