Imagine asking a hotel receptionist for help with a broken air conditioner.
One receptionist can only read the sentence you type. Another can also inspect a photo of the control panel, listen to the strange noise you recorded and read the PDF manual you upload.
That second experience is a useful intuition for multimodal artificial intelligence (Multimodal AI).
What is a modality?
A modality is simply a form in which information arrives. Text is one modality. Images, audio and video are others.
Earlier AI systems were often specialized: one system classified pictures, another transcribed speech and another processed text. Modern systems increasingly accept several forms of information in one interaction.
If the idea of a model is still fuzzy, revisit Lesson 002. A model is the trained core that turns an input into an output.
Multimodal does not just mean several apps glued together
A pipeline can certainly connect separate systems. Speech recognition might first convert audio into text, then send that text to a language model.
A multimodal model goes further: it can represent more than one type of input and reason about them together.
For example, you can upload a photo of your refrigerator and ask, “What can I cook tonight?” The model must use both the visual evidence and the meaning of your question.
That is different from merely returning a list such as “eggs, tomatoes and milk.”
Does AI see an image the way a person does?
Not exactly.
Humans experience space, faces and context directly. A model converts an image into numerical representations that can be processed together with other information.
That means multimodal AI can still:
- miss tiny text,
- confuse the position of objects,
- misread charts,
- invent a detail that is not present,
- misunderstand an ambiguous frame in a video.
The warning from Lesson 006 still applies: an answer can sound reasonable without being factually correct.
The useful part is combining clues
Think of a detective. A written statement is useful, but a photograph, recording, video and document can reveal different parts of the same event.
Multimodal AI becomes valuable when those sources are combined in one task. It can, for example:
- inspect a product photo and draft a description,
- listen to a meeting and extract action items,
- read a chart together with its explanatory text,
- examine a short video and answer when an action happened,
- compare a design mockup with a requirements document.
Input and output are separate questions
A model that can accept images does not automatically generate images.
When evaluating a model, ask two questions:
- What can it receive as input?
- What can it produce as output?
A translator may understand three spoken languages but still be asked to deliver only a written English transcript. AI capabilities work the same way.
One thing to remember
Multimodal AI can work with different forms of information—such as text, images, audio and video—inside the same task.
In Lesson 013, we move from “AI can inspect an image” to a different question: how can AI create an image that did not exist before?
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.