← Back home
LESSON 012AI Basics7 min

What is multimodal AI? Think of an assistant who can read, see and listen

AI does not have to work with text alone. Lesson 012 explains multimodal AI through text, images, audio and video, and why combining clues matters.

Today’s analogyan assistant upgraded from text-only support to seeing images, hearing audio and reading documents

Imagine asking a hotel receptionist for help with a broken air conditioner.

One receptionist can only read the sentence you type. Another can also inspect a photo of the control panel, listen to the strange noise you recorded and read the PDF manual you upload.

That second experience is a useful intuition for multimodal artificial intelligence (Multimodal AI).

What is a modality?

A modality is simply a form in which information arrives. Text is one modality. Images, audio and video are others.

Earlier AI systems were often specialized: one system classified pictures, another transcribed speech and another processed text. Modern systems increasingly accept several forms of information in one interaction.

If the idea of a model is still fuzzy, revisit Lesson 002. A model is the trained core that turns an input into an output.

Multimodal does not just mean several apps glued together

A pipeline can certainly connect separate systems. Speech recognition might first convert audio into text, then send that text to a language model.

A multimodal model goes further: it can represent more than one type of input and reason about them together.

For example, you can upload a photo of your refrigerator and ask, “What can I cook tonight?” The model must use both the visual evidence and the meaning of your question.

That is different from merely returning a list such as “eggs, tomatoes and milk.”

Does AI see an image the way a person does?

Not exactly.

Humans experience space, faces and context directly. A model converts an image into numerical representations that can be processed together with other information.

That means multimodal AI can still:

The warning from Lesson 006 still applies: an answer can sound reasonable without being factually correct.

The useful part is combining clues

Think of a detective. A written statement is useful, but a photograph, recording, video and document can reveal different parts of the same event.

Multimodal AI becomes valuable when those sources are combined in one task. It can, for example:

Input and output are separate questions

A model that can accept images does not automatically generate images.

When evaluating a model, ask two questions:

  1. What can it receive as input?
  2. What can it produce as output?

A translator may understand three spoken languages but still be asked to deliver only a written English transcript. AI capabilities work the same way.

One thing to remember

Multimodal AI can work with different forms of information—such as text, images, audio and video—inside the same task.

In Lesson 013, we move from “AI can inspect an image” to a different question: how can AI create an image that did not exist before?

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. Google — Gemini API Getting Started: Multimodal Understanding ↗
  2. OpenAI — Models ↗
← Previous011How deep can AI roleplay go? Companionship, romance and adult themes depend on more than the model
Next →013How does AI generate images? Think of shaping a noisy canvas into a picture
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200