You may have seen a still illustration suddenly speak, blink and gesture, or a stylized character imitate the expressions of a real performer.
These systems are often grouped under AI avatar or digital human tools, but the underlying workflows can differ significantly.
The simplest avatar combines a face and audio
A basic talking-avatar pipeline can be:
Character image
+ speech audio
→ talking character video
The central operation is often the lip-sync idea from Lesson 040: align the mouth and face with an audio track, then add small head and eye movement.
This works well for explainers, virtual presenters and short educational clips.
If the character must actually perform, however, mouth movement is not enough.
Performance capture supplies the acting
Performance capture records or interprets a human performance and transfers aspects of it to another character.
The source performance may provide:
- facial expressions,
- head direction,
- body movement,
- hand gestures,
- speaking rhythm.
Runway’s Act-Two workflow is one current example: a driving performance video supplies the acting while a character image or video supplies the visual identity.
The driving performance is the acting source
Imagine recording yourself doing this sequence:
frown
→ point right
→ smile
→ step backward
The final character performing those actions could be a robot, cartoon figure or brand mascot.
The source video says how to move.
The character reference says who appears
Lesson 038 explained why references provide a stronger identity anchor than text alone.
In an avatar workflow, it is useful to separate:
Driving performance = acting
Character reference = appearance
That separation is a major reason modern digital-character tools can produce richer performances than a classic talking-head effect.
Gesture control goes beyond the face
A weak talking-head result often looks like a moving face attached to a frozen body.
Fuller performance transfer can incorporate shoulders, hands and posture. The larger the movement, however, the more opportunities there are for occlusion, distorted fingers, clothing drift and body-proportion errors.
Full-body avatar animation is therefore generally harder than a stable chest-up presenter.
Real-time avatars add another requirement: latency
Some avatars are generated offline. Others aim for real-time avatars that respond during a conversation:
User speaks
→ AI understands
→ speech is generated
→ face and gestures animate
→ response appears almost immediately
Here the system must balance quality with low latency, the delay between input and visible response.
Real people introduce identity and consent questions
For fictional characters, the main concerns are asset and platform licenses.
For a real person, also consider:
- likeness rights,
- voice rights,
- consent,
- deepfake risks,
- commercial-use permission.
The fact that a system can combine one person’s face, another voice and a third performance does not automatically grant permission to do so.
A beginner-friendly workflow
- Create one clean, approved character reference.
- Test a short talking-avatar clip.
- Verify voice and lip synchronization.
- Add small gestures.
- Attempt large full-body movement only after the simpler version is stable.
Layering complexity makes failures easier to diagnose.
One thing to remember
Modern AI avatars can separate identity from performance: a character reference determines appearance while a driving performance supplies expressions and movement, with speech and lip sync added as separate layers.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.