← Back home
LESSON 041AI Tools8 min

What are AI avatars and digital humans? A reference image can borrow a performance, voice and expression

AI avatars go beyond moving lips on a photo. Learn how driving performances, character references, gestures and lip sync can be combined to animate a digital character.

Today’s analogyA performer supplies the acting while a character reference acts like a digital costume that determines who appears on screen.

You may have seen a still illustration suddenly speak, blink and gesture, or a stylized character imitate the expressions of a real performer.

These systems are often grouped under AI avatar or digital human tools, but the underlying workflows can differ significantly.

The simplest avatar combines a face and audio

A basic talking-avatar pipeline can be:

Character image
+ speech audio
→ talking character video

The central operation is often the lip-sync idea from Lesson 040: align the mouth and face with an audio track, then add small head and eye movement.

This works well for explainers, virtual presenters and short educational clips.

If the character must actually perform, however, mouth movement is not enough.

Performance capture supplies the acting

Performance capture records or interprets a human performance and transfers aspects of it to another character.

The source performance may provide:

Runway’s Act-Two workflow is one current example: a driving performance video supplies the acting while a character image or video supplies the visual identity.

The driving performance is the acting source

Imagine recording yourself doing this sequence:

frown
→ point right
→ smile
→ step backward

The final character performing those actions could be a robot, cartoon figure or brand mascot.

The source video says how to move.

The character reference says who appears

Lesson 038 explained why references provide a stronger identity anchor than text alone.

In an avatar workflow, it is useful to separate:

Driving performance = acting
Character reference = appearance

That separation is a major reason modern digital-character tools can produce richer performances than a classic talking-head effect.

Gesture control goes beyond the face

A weak talking-head result often looks like a moving face attached to a frozen body.

Fuller performance transfer can incorporate shoulders, hands and posture. The larger the movement, however, the more opportunities there are for occlusion, distorted fingers, clothing drift and body-proportion errors.

Full-body avatar animation is therefore generally harder than a stable chest-up presenter.

Real-time avatars add another requirement: latency

Some avatars are generated offline. Others aim for real-time avatars that respond during a conversation:

User speaks
→ AI understands
→ speech is generated
→ face and gestures animate
→ response appears almost immediately

Here the system must balance quality with low latency, the delay between input and visible response.

For fictional characters, the main concerns are asset and platform licenses.

For a real person, also consider:

The fact that a system can combine one person’s face, another voice and a third performance does not automatically grant permission to do so.

A beginner-friendly workflow

  1. Create one clean, approved character reference.
  2. Test a short talking-avatar clip.
  3. Verify voice and lip synchronization.
  4. Add small gestures.
  5. Attempt large full-body movement only after the simpler version is stable.

Layering complexity makes failures easier to diagnose.

One thing to remember

Modern AI avatars can separate identity from performance: a character reference determines appearance while a driving performance supplies expressions and movement, with speech and lip sync added as separate layers.

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. Runway — Performance Capture with Act-Two ↗
  2. Runway Academy — Character Animation with Act-Two ↗
  3. Runway — Creating with Gen-4 Image References ↗
← Previous040AI video can generate sound now—but native audio, lip sync, TTS and sound effects are different jobs
Next →042A practical AI video workflow: script → storyboard → references → shots → editing, not one giant prompt
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200