Foundations

Multimodal AI

Multimodal model

AI that can take in, and sometimes produce, more than one kind of content, such as text, images, audio and video, within a single model.

textimagesaudiovideo and documentsMultimodal modelone shared representationOutputtext, images, audioOne model can read a chart, listen to a meeting and reason about the two together.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/multimodal-ai

In plain terms

People do not experience the world as text alone; we read, look and listen at once. A multimodal model does something similar. You can show it a photo of a whiteboard, a chart from a report or a recording of a meeting and ask about it in plain language, and it reasons across what it sees, hears and reads.

Why it matters

Most business information is not clean text. It is scanned invoices, slides, diagrams, photographs from site inspections, screenshots and phone calls. Multimodal models make this material usable without a separate system for each format. They are also what allows agents to operate software by looking at the screen. Accuracy depends on the task: reading a clear chart is reliable; counting small objects or reading fine detail in a poor scan is less so.

Example

A field technician photographs a damaged electrical cabinet and asks aloud what the fault is likely to be. The model looks at the photo, reads the model number on the label, hears the question, and answers with the probable cause and the page of the service manual to consult. Three kinds of input, one model, one answer.

Most often confused with

Multimodal AI vs. Separate models chained together

Multimodal AIOne model reasons across all inputs together
Separate models chained togetherEach input is turned into text first, then passed on

Before multimodal models, a system that handled speech would transcribe it with one model and pass the text to another. That works, but information is lost at every handover: the tone of a voice, the layout of a page, where an element sits in an image. A natively multimodal model keeps that information and can relate it across inputs.

Under the hood

Each kind of input is converted into tokens or embeddings in a shared representation: images are cut into patches and encoded, audio into short frames. The model's attention then operates over all of them together. Input and output types differ by model: many accept text, images, audio and video and produce text; some also generate images or speech natively. An earlier building block is CLIP, which aligned images and text in one embedding space. Practical points: images and video consume many tokens, and so cost and context; resolution limits affect small text and fine detail; documents are often best supplied as page images together with extracted text. Uses: document understanding, questions about images, voice assistants, video analysis, computer use. A risk to note: instructions hidden in images or audio are a route for prompt injection.

Written by Mehmet Erkek · Last updated: