← Back home
LESSON 037AI Tools9 min

How should you generate AI video in 2026? Start with text, an image, references or existing video

AI video is no longer one prompt box. Learn when to use text-to-video, image-to-video, references, first/last frames and video-to-video before choosing a model.

Today’s analogyChoosing a video mode is like deciding whether a director starts from a script, a key image or a full reference pack.

Open an AI video product today and the first thing you may see is a long list of model names.

For a beginner, however, the more useful first question is not “Which model is strongest?” It is “What am I giving the model as its starting point?”

Lesson 016 introduced the central difficulty of video generation: a model must generate not only a good frame, but a sequence that remains coherent over time.

Modern video tools expose several ways to give that sequence a starting structure.

Text-to-video starts from an idea

Text-to-video (T2V) means the model receives a text description and invents the scene from scratch.

For example:

A rainy Taipei alley at night. A person in a red coat walks past neon signs.
Low-angle tracking shot, reflections on wet pavement, cinematic lighting, 6 seconds.

This is excellent for exploration because the model has a great deal of freedom. You can test mood, composition, action and camera language quickly.

The same freedom is also the weakness: the model must decide the face, clothing details, object placement and many other visual facts on its own. Those details can drift between attempts.

Use text-to-video when the idea matters more than preserving an exact visual identity.

Image-to-video fixes the visual starting point

Image-to-video (I2V) begins with an image and asks the model to animate it.

This connects directly to Lesson 013: you can first design a character, product or scene as a still image, approve it, and then turn that approved frame into motion.

A useful prompt becomes more focused:

She slowly turns toward the window and rotates the glass in her right hand.
The camera makes a very gentle push-in. Keep the composition stable.

You no longer need to spend most of the prompt re-describing what the person looks like. The image carries much of that information.

References give the model more than one clue

Many current workflows also accept references: images or videos that specify identity, wardrobe, visual style, motion, scene design or camera behavior.

Think of a production team receiving not only a storyboard but also a casting photo, costume reference and movement example.

The exact reference controls vary by product and model version, so treat “reference support” as a capability to inspect rather than a universal checkbox.

First and last frames constrain the transition

Some systems accept a first frame and last frame.

Instead of asking for an unconstrained transformation, you can define:

Start: an intact white rose on a table
End: the rose has transformed into scattered glass fragments

The task becomes: create a plausible path from state A to state B.

This is useful for transitions, transformations, product reveals and shots where the ending composition matters.

Video-to-video keeps an existing motion skeleton

Video-to-video (V2V) starts from an existing clip. Depending on the tool, it can preserve parts of the original motion, timing or camera path while changing the appearance, character or style.

That makes it useful when you already have a performance or camera move that works and do not want the model to reinvent it.

Current systems increasingly combine input types

As of September 2026, first-party documentation for platforms such as Google Veo 3.1 and Runway shows that modern video generation is not limited to one text-only workflow. Different model versions expose text, images, references, existing video, frame constraints and extension/editing capabilities.

Capabilities change quickly, so a comparison table from six months ago may already be stale. The transferable decision is the workflow:

Exploring from scratch?
→ Text-to-video

Already have the visual you want?
→ Image-to-video

Need identity, style or motion continuity?
→ References / video-to-video

Care strongly about the opening and ending?
→ First / last frames

The best demo model may not be the best production model

Suppose your real goal is eight shots of the same character holding the same product. In that case, reference control may matter more than whether one text-to-video benchmark clip looks spectacular.

When evaluating a video model, ask:

  1. What input types does it accept?
  2. Can it preserve references?
  3. Can you constrain first or last frames?
  4. Can it extend or edit existing footage?
  5. Is audio generated with the video or separately?
  6. Can the output fit your editing workflow?

One thing to remember

AI video is no longer one “prompt to clip” problem. Choose the input mode that matches what you already know and what you need to control; only then compare models.

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. Google Cloud — Veo 3.1 ↗
  2. Runway Developer — Available AI Models ↗
  3. Runway Developer — API Input Parameters ↗
← Previous036Guardrails and “uncensored” models: safety is a stack, not one hidden switch
Next →038Why does the same person keep changing in AI video? Character consistency needs more than a prompt
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200