A single AI image only has to look convincing at one moment.
Video has a second requirement: the next moment must make sense too.
If a person turns their head, their face should remain the same person. If a cup is on a table, it should not disappear randomly. If the camera moves left, the room should reveal new space coherently.
That extra dimension is time.
Think of a flipbook
A flipbook works because every page is slightly different from the previous one. Each drawing must be good enough on its own, but the sequence also has to create believable motion.
AI video generation faces the same two-level problem:
- each frame must look plausible,
- neighboring frames must remain consistent with one another.
This second requirement is often called temporal consistency.
Image generation is part of the problem, not the whole problem
The intuition from Lesson 013 still helps. Video models learn visual patterns and generate frames under conditions such as text, reference images or an initial frame.
But now they must also model movement and change over time.
Depending on the system, inputs may include:
- a text prompt,
- a starting image,
- first and last frames,
- a reference character,
- an existing video to transform,
- camera or motion controls.
Why do AI videos melt or morph?
If the model loses track of an object or identity across time, small errors compound.
A hand may change shape, a necklace may vanish, text may mutate or a background object may move without a physical reason.
These are not merely “bad frames.” They are failures to maintain a stable world model through time.
Motion has several layers
When you ask “the woman walks toward the camera,” at least three things can move:
- the subject,
- the environment,
- the camera.
A useful video prompt separates them.
For example: “The subject walks slowly forward; her coat moves lightly in the wind; the camera tracks backward at the same pace; the city background remains stable.”
That is clearer than simply writing “cinematic walking shot.”
Duration matters
A five-second clip is not just a smaller version of a one-minute film. Longer sequences give more time for identities, objects and geometry to drift.
In real production, creators often generate short shots and edit them together rather than demanding a full scene from one generation.
This also matches normal filmmaking: a sequence is built from shots.
Reference images become even more important
The consistency ideas from Lesson 015 apply strongly to video. If a character must remain recognizable, reference images or identity controls can provide a stable anchor.
Without them, “same woman” may be interpreted too loosely over many frames.
One thing to remember
AI video generation must create plausible images and plausible change between those images. Time is the extra difficulty.
Lesson 017 turns this into a practical writing method for video prompts: subject motion, camera motion, scene and timing.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.