Open an AI video product today and the first thing you may see is a long list of model names.
For a beginner, however, the more useful first question is not “Which model is strongest?” It is “What am I giving the model as its starting point?”
Lesson 016 introduced the central difficulty of video generation: a model must generate not only a good frame, but a sequence that remains coherent over time.
Modern video tools expose several ways to give that sequence a starting structure.
Text-to-video starts from an idea
Text-to-video (T2V) means the model receives a text description and invents the scene from scratch.
For example:
A rainy Taipei alley at night. A person in a red coat walks past neon signs.
Low-angle tracking shot, reflections on wet pavement, cinematic lighting, 6 seconds.
This is excellent for exploration because the model has a great deal of freedom. You can test mood, composition, action and camera language quickly.
The same freedom is also the weakness: the model must decide the face, clothing details, object placement and many other visual facts on its own. Those details can drift between attempts.
Use text-to-video when the idea matters more than preserving an exact visual identity.
Image-to-video fixes the visual starting point
Image-to-video (I2V) begins with an image and asks the model to animate it.
This connects directly to Lesson 013: you can first design a character, product or scene as a still image, approve it, and then turn that approved frame into motion.
A useful prompt becomes more focused:
She slowly turns toward the window and rotates the glass in her right hand.
The camera makes a very gentle push-in. Keep the composition stable.
You no longer need to spend most of the prompt re-describing what the person looks like. The image carries much of that information.
References give the model more than one clue
Many current workflows also accept references: images or videos that specify identity, wardrobe, visual style, motion, scene design or camera behavior.
Think of a production team receiving not only a storyboard but also a casting photo, costume reference and movement example.
The exact reference controls vary by product and model version, so treat “reference support” as a capability to inspect rather than a universal checkbox.
First and last frames constrain the transition
Some systems accept a first frame and last frame.
Instead of asking for an unconstrained transformation, you can define:
Start: an intact white rose on a table
End: the rose has transformed into scattered glass fragments
The task becomes: create a plausible path from state A to state B.
This is useful for transitions, transformations, product reveals and shots where the ending composition matters.
Video-to-video keeps an existing motion skeleton
Video-to-video (V2V) starts from an existing clip. Depending on the tool, it can preserve parts of the original motion, timing or camera path while changing the appearance, character or style.
That makes it useful when you already have a performance or camera move that works and do not want the model to reinvent it.
Current systems increasingly combine input types
As of September 2026, first-party documentation for platforms such as Google Veo 3.1 and Runway shows that modern video generation is not limited to one text-only workflow. Different model versions expose text, images, references, existing video, frame constraints and extension/editing capabilities.
Capabilities change quickly, so a comparison table from six months ago may already be stale. The transferable decision is the workflow:
Exploring from scratch?
→ Text-to-video
Already have the visual you want?
→ Image-to-video
Need identity, style or motion continuity?
→ References / video-to-video
Care strongly about the opening and ending?
→ First / last frames
The best demo model may not be the best production model
Suppose your real goal is eight shots of the same character holding the same product. In that case, reference control may matter more than whether one text-to-video benchmark clip looks spectacular.
When evaluating a video model, ask:
- What input types does it accept?
- Can it preserve references?
- Can you constrain first or last frames?
- Can it extend or edit existing footage?
- Is audio generated with the video or separately?
- Can the output fit your editing workflow?
One thing to remember
AI video is no longer one “prompt to clip” problem. Choose the input mode that matches what you already know and what you need to control; only then compare models.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.