A still-image prompt can describe what a frame should look like. A video prompt must also describe what changes after the first frame.
If you simply write “a cinematic woman walking through a futuristic city,” the model still has to guess whether she walks toward the camera, across the frame or away from it, and whether the camera is static, tracking or orbiting.
A useful video prompt therefore behaves more like a short shot description.
Start with one shot, not an entire movie
Video generators are usually more reliable when one prompt describes one coherent shot.
Instead of asking for “a woman leaves home, takes a taxi, arrives at a hotel, enters the lobby and meets a friend,” split the sequence into several clips.
This reduces the number of identities, locations and transitions the model must keep consistent at once.
Use five parts
A reusable structure is:
- Subject — who or what is in the shot?
- Subject motion — what exactly does it do?
- Camera motion — what does the camera do?
- Scene and lighting — where is it, and what is the visual mood?
- Timing — how fast or slow should the action feel?
For example:
A woman in a long black coat walks slowly toward the camera through a wet neon alley. Her coat and hair move lightly in the wind. The camera tracks backward smoothly at walking speed. Reflections remain stable on the pavement. Night, soft blue and magenta practical lights, restrained cinematic motion.
Notice that the subject and camera are described separately.
Describe motion with verbs
“Elegant” is a style. “Turns her head slowly, pauses, then looks back” is motion.
Useful motion language includes:
- walks forward,
- turns left,
- reaches for the cup,
- raises one hand,
- fabric moves in the wind,
- camera pans right,
- camera tracks backward,
- slow push-in,
- static locked-off camera.
The goal is to reduce ambiguity about change over time.
Do not overload a five-second clip
If a short generation contains six major actions, the model may compress, skip or blend them.
For short clips, one main action plus one camera movement is often easier to control.
Think like an editor: if the story needs a new action or viewpoint, that may deserve a new shot.
Separate camera movement from subject movement
“Camera moves left while the subject walks right” is much clearer than “dynamic movement.”
This distinction also helps diagnose failures. If the person moves correctly but the framing drifts, revise the camera instruction rather than changing the character description.
Keep visual anchors from earlier lessons
The prompt structure from Lesson 014 still matters: subject, environment, composition and lighting.
The consistency methods from Lesson 015 matter even more when the same character appears in several shots.
And the temporal-consistency problem from Lesson 016 explains why a simple prompt can still produce morphing details.
One thing to remember
A good video prompt describes a shot: what moves, what the camera does, where the scene is and how the action unfolds over time.
From Lesson 018 onward, we leave media generation and enter a more technical path: how applications actually talk to AI models through APIs.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.