Home › The Road to LLMs › Word embeddings
🛤️ · Foundations

Word embeddings

Dense numeric vectors that place words with similar meaning near each other in space.

In one line

Word embeddings turn words into vectors so machines can measure meaning as distance.

ConceptWhat it is

Word embeddings are dense numeric vectors, typically a few hundred dimensions, that represent words such that semantically similar words sit close together in vector space. Methods like Word2Vec and GloVe popularized them around 2013, replacing sparse one-hot word representations.

They exist because raw text has no inherent numeric structure a model can compute over; embeddings give language a geometry where meaning becomes measurable, enabling arithmetic-like relationships such as king minus man plus woman approximating queen.

How it worksThe mechanics

A model is trained to predict a word from its surrounding context, or vice versa, over a huge text corpus; the byproduct of that training is a lookup table of vectors where words appearing in similar contexts end up with similar vectors.

At a glanceSee it

Word embeddings diagram
Word embeddings diagram 1

Because related word pairs share a consistent direction in the space, plain vector arithmetic can answer analogies — king minus man plus woman lands nearest to queen.

Word embeddings diagram 2

The two classic recipes split on whether they mine global co-occurrence counts or local prediction windows, yet both land on the same dense-vector format that later models consume.

When to use itWhere it fits

  • Search and recommendation systems needing semantic similarity, not just keyword match.
  • Feature input for downstream classical or deep NLP models.
  • Clustering or visualizing large text corpora by topic.
  • Bootstrapping simple semantic search before adopting full transformer embeddings.

When NOT to use itLimits & anti-patterns

  • Tasks needing full sentence or document-level context, since single-word embeddings ignore word order.
  • Highly polysemous language where a single static vector per word is too coarse.
  • Modern production semantic search, where transformer-based embeddings now outperform static ones.

Trade-offsAdvantages & costs

Advantages
  • Compact, fast to compute similarity over.
  • Captures meaningful semantic relationships.
  • Cheap to train and use compared to full language models.
  • Easy to plug into classical ML pipelines.
Trade-offs & costs
  • One vector per word regardless of context, so it cannot distinguish river bank from financial bank.
  • Superseded in most modern systems by contextual transformer embeddings.
  • Requires a large corpus to train good vectors.
  • Struggles with rare words and out-of-vocabulary terms.

ExampleIn the real world

Early Google News recommendation systems used Word2Vec-trained vectors to cluster and surface semantically related articles before transformer embeddings existed.

ToolsHow to implement it

  • Word2Vecthe original efficient context-prediction embedding method.
  • GloVeco-occurrence-statistics-based alternative from Stanford.
  • fastTextFacebook's subword-aware embeddings, robust to rare words.
  • Gensimpopular Python library for training and using word vectors.

Cost & effortWhat it takes

Very cheap: training takes minutes to hours on a CPU or single GPU over a large corpus; lookup at inference is essentially free; low engineering effort versus modern contextual embeddings.

A living map of modern AI — kept current every morning