Home › Embeddings & Vector Search › Multimodal (CLIP)
🧭 · Models

Multimodal (CLIP)

Mapping images and text into one shared vector space for cross-modal search.

In one line

CLIP-style models let you search images with text and text with images, in one space.

ConceptWhat it is

Multimodal embeddings, popularized by OpenAI's CLIP, map images and text into the same vector space so a photo of a dog and the word dog land near each other. They exist because unimodal systems cannot answer "find the image that matches this caption" without a shared representation.

Training uses millions of image-caption pairs with a contrastive loss, teaching the model that matched pairs should be close and mismatched pairs far apart, without needing manual labels.

How it worksThe mechanics

An image encoder and a text encoder are trained jointly; each produces a vector in the same dimensional space, and similarity between any image vector and any text vector is computed with cosine distance, enabling text-to-image search, image-to-image search, and zero-shot image classification by comparing an image to a set of label embeddings.

At a glanceSee it

Multimodal (CLIP) diagram
Multimodal (CLIP) diagram 1

The contrastive training loop CLIP repeats over millions of batches — score every image against every caption in the batch, then pull matched pairs together and push all mismatches apart.

Multimodal (CLIP) diagram 2

How the same shared space does zero-shot classification without training any classifier — label names become text prompts, the nearest one wins, and a threshold lets it abstain when the answer is not among the candidates.

When to use itWhere it fits

  • Visual search in e-commerce, like find shoes similar to this photo.
  • Zero-shot image tagging without training a classifier.
  • Content moderation matching images against text policy descriptions.
  • Cross-modal recommendation, pairing product images with text queries.

When NOT to use itLimits & anti-patterns

  • Fine-grained classification tasks where a supervised classifier outperforms zero-shot CLIP.
  • Precise OCR or reading text within images, which CLIP handles poorly.
  • Domains far from web-scale training data, like medical imaging, without fine-tuning.

Trade-offsAdvantages & costs

Advantages
  • Enables true cross-modal search in one embedding space.
  • Zero-shot classification without labeled training data.
  • Transfers well across many visual domains.
  • Simple to combine with existing vector databases.
Trade-offs & costs
  • Weak at fine-grained detail and counting objects.
  • Inherits biases from web-scraped training data.
  • Larger compute cost than text-only embeddings.
  • Struggles outside natural-image domains without fine-tuning.

ExampleIn the real world

Pinterest uses CLIP-style multimodal embeddings to power visual search, letting users screenshot an object and find visually and semantically similar pins.

ToolsHow to implement it

  • OpenAI CLIPthe original open-weights image-text encoder.
  • OpenCLIPopen-source reproductions trained on larger datasets.
  • Google SigLIPimproved contrastive loss for better zero-shot accuracy.
  • Weaviatevector database with native multimodal module support.

Cost & effortWhat it takes

Moderate compute cost for image encoding, higher than text-only embeddings; latency in the low hundreds of milliseconds per image on GPU.

A living map of modern AI — kept current every morning