Contact

What is Multimodal Model?

Definition

A multimodal model is an AI model that can take more than one type of data, or modality, as input or produce it as output: text, images, audio or video. Such a model can answer questions about a photo, read a chart or generate an image from a description. Because the different modalities are processed in a shared representation space, the model can connect an object in an image with the words that describe it.

Also known as: multimodal AI, multimodal LLM, vision-language model, VLM, MLLM

Diagram of a multimodal model taking and producing text, images, video, charts, PDFs and code in one model

Turning pixels into tokens

Language models work on sequences of tokens, so the trick in a multimodal model is to turn other data into something that can join that sequence. An image is typically cut into small patches; an image encoder converts each patch into a vector, and those vectors are projected into the same space as the text tokens. Audio is handled similarly, in short time slices. From there, the Transformer layers can attend across image patches and words alike, which is how a question like “what does the label in the bottom-left corner say?” gets answered.

The same shared-space idea underlies embedding models that place images and text in one vector space, which is what makes searching a photo library with a sentence possible.

Common combinations

Input → outputExample uses
Image + text → textExplaining an error in a screenshot, pulling fields from an invoice, reading a chart
Text → imageDraft illustrations, variations on a campaign visual
Audio → textMeeting transcripts, voice commands
Text → audioVoice-over, spoken assistant replies
Video → textSummarising a recording, finding a specific moment

No single model necessarily does all of these. Some only understand images and reply in text; others can also generate images or speech. Which modalities a model supports is stated in its model card or the provider's documentation, and it changes between versions.

What changes for web content

AI systems can now interpret screenshots, PDFs, tables and product photos. Information can therefore be read from formats other than text, but that does not make text less important:

  • Most search and retrieval systems still index text first. A price or offer that exists only inside a banner image is less likely to be found.
  • Alt text is an accessibility requirement for people using screen readers. A model's ability to describe an image on its own does not remove that obligation.
  • File names, captions and nearby descriptive text help image SEO and help multimodal systems place an image in the right context.

Known weak spots

  • Visual hallucination: describing objects that aren't there, or misreading small print.
  • Counting and spatial reasoning: how many items, which is left of which, and relative sizes are frequent failure points.
  • Privacy: uploading an invoice, an ID or a screenshot also sends any personal data in it. Teams using these tools internally should be clear about what goes to which provider.

Many recent foundation models are trained as multimodal from the start, so the line between a “language model” and a “vision model” is getting harder to draw.

Related terms

← Back to the glossary