← Back home
LESSON 017AI Tools9 min

How to write AI video prompts: separate subject motion, camera motion, scene and timing

Video prompts need more than a pretty scene. Lesson 017 shows how to describe subject motion, camera movement, environment, timing and shot boundaries clearly.

Today’s analogygiving a film crew a short shot list instead of saying only “make it cinematic”

A still-image prompt can describe what a frame should look like. A video prompt must also describe what changes after the first frame.

If you simply write “a cinematic woman walking through a futuristic city,” the model still has to guess whether she walks toward the camera, across the frame or away from it, and whether the camera is static, tracking or orbiting.

A useful video prompt therefore behaves more like a short shot description.

Start with one shot, not an entire movie

Video generators are usually more reliable when one prompt describes one coherent shot.

Instead of asking for “a woman leaves home, takes a taxi, arrives at a hotel, enters the lobby and meets a friend,” split the sequence into several clips.

This reduces the number of identities, locations and transitions the model must keep consistent at once.

Use five parts

A reusable structure is:

  1. Subject — who or what is in the shot?
  2. Subject motion — what exactly does it do?
  3. Camera motion — what does the camera do?
  4. Scene and lighting — where is it, and what is the visual mood?
  5. Timing — how fast or slow should the action feel?

For example:

A woman in a long black coat walks slowly toward the camera through a wet neon alley. Her coat and hair move lightly in the wind. The camera tracks backward smoothly at walking speed. Reflections remain stable on the pavement. Night, soft blue and magenta practical lights, restrained cinematic motion.

Notice that the subject and camera are described separately.

Describe motion with verbs

“Elegant” is a style. “Turns her head slowly, pauses, then looks back” is motion.

Useful motion language includes:

The goal is to reduce ambiguity about change over time.

Do not overload a five-second clip

If a short generation contains six major actions, the model may compress, skip or blend them.

For short clips, one main action plus one camera movement is often easier to control.

Think like an editor: if the story needs a new action or viewpoint, that may deserve a new shot.

Separate camera movement from subject movement

“Camera moves left while the subject walks right” is much clearer than “dynamic movement.”

This distinction also helps diagnose failures. If the person moves correctly but the framing drifts, revise the camera instruction rather than changing the character description.

Keep visual anchors from earlier lessons

The prompt structure from Lesson 014 still matters: subject, environment, composition and lighting.

The consistency methods from Lesson 015 matter even more when the same character appears in several shots.

And the temporal-consistency problem from Lesson 016 explains why a simple prompt can still produce morphing details.

One thing to remember

A good video prompt describes a shot: what moves, what the camera does, where the scene is and how the action unfolds over time.

From Lesson 018 onward, we leave media generation and enter a more technical path: how applications actually talk to AI models through APIs.

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. Google — Video generation with Veo in the Gemini API ↗
  2. OpenAI — Video generation API ↗
← Previous016How does AI generate video? It has to keep a moving world consistent across time
Next →018What is an API? Think of a restaurant order window between your app and an AI model
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200