Home › Embeddings & Vector Search › Cohere Embed v4
🧭 · Models

Cohere Embed v4

Cohere Embed v4: enterprise multilingual and multimodal embeddings with adjustable Matryoshka dimensions and long context

In one line

A hosted embedding model, accessed through an API, that turns text and images in 100-plus languages into vectors you can shrink from 1536 down to 256 dimensions without re-encoding.

ConceptWhat it is

Cohere Embed v4 is a hosted, proprietary text and image embedding model that maps documents and queries into dense vectors for semantic search and retrieval-augmented generation (RAG). It is built for enterprise workloads: it understands 100-plus languages, accepts inputs up to a 128K-token context, and can embed mixed text-and-image content — a slide, a scanned form, a PDF page with tables and charts — into a single searchable vector.

Its defining trait is Matryoshka representation learning: one API call returns a vector whose leading dimensions carry the most information, so you can truncate from 1536 down to 256 dimensions and still retrieve well, trading a little accuracy for large savings in storage and search. Embed v4 also emits compressed formats such as int8 and binary, letting teams tune the memory footprint of a very large index. Because it is a hosted, managed model, you call an endpoint and get vectors back rather than downloading and serving the weights yourself.

How it worksThe mechanics

You send a batch of texts or images to the embed endpoint with an input_type hint that tells the model whether the content is a document to index or a query to search with, plus an optional target dimension and output format. The model tokenizes the input, runs it through its transformer encoder, and pools the result into one dense vector per item; for a mixed page it fuses the visual and textual signal into a single vector. You store document vectors in a vector index, then at query time embed the question the same way and run a nearest-neighbor search — cosine or dot-product similarity — to pull the top matches, which you feed to a generator or rerank before returning the final answer.

At a glanceSee it

Cohere Embed v4 diagram

When to use itWhere it fits

  • Building multilingual RAG where a query in one language must retrieve documents written in another.
  • Indexing visually rich content — PDFs, slides, screenshots, charts, scanned forms — without a separate OCR and parsing pipeline.
  • Operating at large scale where truncating dimensions or using int8 and binary vectors materially cuts storage and query cost.
  • Teams that want top-tier retrieval quality without hosting, patching, or scaling their own embedding infrastructure.

When NOT to use itLimits & anti-patterns

  • Air-gapped or strict data-residency settings that forbid sending content to a third-party API, unless a permitted private cloud deployment is available.
  • Cost-sensitive, very high-volume pipelines where a free self-hosted open model already meets the quality bar.
  • Cases needing a fully inspectable or deeply fine-tuned embedding model, since the weights are closed and served only through the vendor.
  • Tiny corpora or exact-match lookups where keyword search or a lightweight local model is simpler and cheaper.

Trade-offsAdvantages & costs

Advantages
  • Strong out-of-the-box retrieval quality across 100-plus languages, reducing per-language tuning.
  • Native multimodal embeddings collapse text and image search into one vector space and one index.
  • Matryoshka dimensions plus int8 and binary output give a direct dial between accuracy and storage or latency cost.
  • Long 128K context lets whole documents be embedded without aggressive chunking gymnastics.
Trade-offs & costs
  • Recurring per-token API cost that scales with corpus size and every re-embed, unlike a one-time self-hosted setup.
  • Vendor dependency: no open weights to own, no fully independent offline copy, and no full fine-tuning of the model.
  • Data must leave your environment for the API, raising governance and residency questions.
  • Changing models or major settings forces a full corpus re-embed, which is slow and costly at scale.

ExampleIn the real world

A global insurer builds a claims-support assistant over policy documents, adjuster photos, and scanned incident forms spanning a dozen languages. Each document, including image-heavy pages, is sent to Embed v4 with input_type set to document and stored as int8 vectors truncated to 512 dimensions in a managed vector database, keeping a very large index affordable. When a French-speaking agent asks a question, the query is embedded with the search input_type, nearest-neighbor search returns the most relevant passages and photos regardless of source language, and those results are passed to a chat model to draft a grounded answer with citations. Promoting a high-stakes subset to full 1536 dimensions for extra recall is a config change, not a re-architecture.

ToolsHow to implement it

  • Cohere SDK and REST API (the co.embed call) for native access to Embed v4.
  • Embed v4 hosted on Amazon Bedrock, Amazon SageMaker, and Azure AI Foundry for in-cloud, privately networked deployments.
  • Vector stores such as pgvector, Pinecone, Weaviate, or Qdrant to index and search the vectors.
  • RAG frameworks like LangChain and LlamaIndex, which ship Cohere embedding integrations.

Cost & effortWhat it takes

Pricing is usage-based per token through the hosted API, so indexing cost scales with total corpus tokens and every re-embed repeats it; query embeddings are cheap individually but add up at high query rates. The real levers sit downstream: dimension count and output format. A 256-dimension int8 vector can cut vector-store cost by roughly an order of magnitude versus full 1536-dimension floats, so most of the cost engineering happens in storage and search rather than in the embedding call itself. Effort to start is low — a single API call — but budget for a full re-embedding pass whenever you change the model or the target dimensions.

A living map of modern AI — kept current every morning