Home › LLMs & Foundation Models › Multimodal
🤖 · Models

Multimodal

Models that natively understand and generate across text, images, audio, and video.

In one line

Multimodal models process and reason across text, images, audio, and video within a single model.

ConceptWhat it is

A multimodal model accepts and often generates more than one type of content, such as text, images, audio, or video, within a unified architecture, rather than stitching together separate specialist models. It exists because so much real-world information is not pure text: a chart, a photo, a voice memo, or a video all carry meaning that a text-only model cannot access directly.

Modern multimodal LLMs typically encode each modality into a shared representation space so the same transformer backbone can reason jointly across, for example, an image and the question asked about it.

How it worksThe mechanics

Non-text inputs like images or audio are passed through a modality-specific encoder that converts them into embeddings compatible with the model's token representation space, which are then interleaved with text token embeddings and processed jointly by the transformer's attention layers; generation can similarly produce text, and in some models images or audio, as output tokens.

At a glanceSee it

Multimodal diagram
Multimodal diagram 1

Zooming inside the modality encoder — an image is patchified, embedded, then projected into the same token space the language model already reads.

Multimodal diagram 2

A routing choice — hand precise extraction to a specialist tool, or let the model reason natively and accept some risk of invented detail.

When to use itWhere it fits

  • Visual question answering, document and chart understanding, and OCR-heavy workflows.
  • Voice assistants that need to process spoken input and respond naturally.
  • Video understanding tasks like summarizing or searching footage.
  • Any product where users naturally provide mixed media, like screenshots plus text questions.

When NOT to use itLimits & anti-patterns

  • Pure text tasks, where a text-only model is cheaper and just as accurate.
  • High-precision specialist vision tasks like medical image segmentation, where a dedicated CNN or vision model may outperform a general multimodal LLM.
  • Extremely cost-sensitive, high-volume workloads where multimodal encoders add unnecessary overhead.

Trade-offsAdvantages & costs

Advantages
  • Single model handles diverse input types without a separate pipeline per modality.
  • Enables richer applications like visual search and voice-first assistants.
  • Joint reasoning across modalities often outperforms bolting together separate models.
  • Rapidly improving quality across image, audio, and video understanding.
Trade-offs & costs
  • Higher inference cost and latency than text-only equivalents.
  • Non-text modalities can still be misread or hallucinated, like misreading chart values.
  • Training requires large, well-aligned multimodal datasets that are harder to source.
  • Evaluation is less mature than for text-only benchmarks.

ExampleIn the real world

Google's Gemini and OpenAI's GPT-4o and GPT-5 natively process text, images, and audio in a single model, powering features like real-time visual assistance in the Gemini and ChatGPT mobile apps.

ToolsHow to implement it

  • GPT-4o / GPT-5 visionAPI access to native multimodal understanding.
  • Google Geminimultimodal model spanning text, image, audio, and video.
  • LLaVA / Qwen-VLopen-weight multimodal model families.
  • Whisperspeech-to-text component often paired into multimodal pipelines.

Cost & effortWhat it takes

Higher per-request cost than text-only, often billed by image resolution or audio duration in addition to tokens; needs paired multimodal training data to build from scratch; latency depends on modality and input size.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Google now offers speech-to-speech models alongside OpenAI's, giving voice-interface builders a second vendor to price and benchmark against.

    Google DeepMind · 15 Sep 2026 · source

  • Updated this page A new experimental variant extends DeepSeek's V4-Flash line to image understanding, giving multimodal builders another model to evaluate.

    Add a note to the DeepSeek V4 Flash page that the line now includes an experimental multimodal vision variant, V4-Flash-Vision-Exp, for image understanding.

    China frontier labs · 23 Aug 2026 · source

  • Updated this page A new benchmark tests multimodal models' abstract perceptual reasoning from dynamic processes, exposing a capability gap beyond static recognition.

    arXiv cs.AI · 17 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning