Text embeddings turn language into coordinates so similar meaning becomes nearby points.
ConceptWhat it is
Text embeddings are dense numeric vectors, typically 384 to 3072 dimensions, produced by a neural encoder so that semantically similar text lands close together in vector space. They exist because keyword matching fails when two sentences mean the same thing but share no words, like "cancel my order" and "I want a refund".
Modern embedding models are trained with contrastive objectives, pulling matching pairs together and pushing unrelated pairs apart, so distance in the vector space becomes a proxy for semantic similarity.
How it worksThe mechanics
Text is tokenized, passed through a transformer encoder, and pooled into a single fixed-length vector, usually via mean pooling or a dedicated CLS token; that vector is then normalized so cosine similarity and dot product behave consistently, and stored or compared against other vectors using distance metrics like cosine or dot product.
At a glanceSee it
How the geometry is actually built — contrastive training pulls matching pairs together and pushes mismatches apart, one batch at a time, until meaning becomes distance.
How one pairwise comparison scales into search — documents are embedded offline into an index, then each query is embedded live and routed to an exact scan or an approximate nearest-neighbor lookup by corpus size.
When to use itWhere it fits
- Building semantic search over documents, support tickets, or product catalogs.
- Powering RAG retrieval so an LLM gets relevant context.
- Deduplication and clustering of large text collections.
- Recommendation systems based on content similarity.
When NOT to use itLimits & anti-patterns
- Exact keyword or regex matching, like legal citation lookup, where embeddings blur precise matches.
- Tiny datasets under a few hundred items, where simple TF-IDF or SQL LIKE is cheaper and just as accurate.
- Real-time low-latency paths where an extra encoder call adds unacceptable delay.
Trade-offsAdvantages & costs
Advantages
- Captures meaning beyond exact word overlap.
- Works across languages with multilingual models.
- Cheap to compute at scale with batching.
- Composable with vector databases and ANN indexes.
Trade-offs & costs
- Opaque; hard to explain why two items matched.
- Quality depends heavily on the embedding model chosen.
- Needs re-embedding when the model is upgraded.
- Struggles with negation and precise numeric facts.
ExampleIn the real world
Notion's search feature embeds every page and block so a query like "Q3 budget notes" surfaces relevant pages even when they never use the word budget verbatim.
ToolsHow to implement it
- OpenAI text-embedding-3strong general-purpose default with tunable dimensions.
- Cohere Embed v3strong multilingual and retrieval-tuned variant.
- sentence-transformersopen-source, self-hostable encoder models.
- Voyage AIdomain-tuned embeddings for code and finance.
Cost & effortWhat it takes
Low cost per call, cents per million tokens, sub-100ms latency; main effort is choosing dimensionality and re-embedding pipeline when models change.