🧠 · Foundations

Encoder (BERT)

A bidirectional transformer built to understand text deeply, not generate it.

In one line

Encoder-only models like BERT read the whole sentence at once to build rich understanding, not to write new text.

ConceptWhat it is

An encoder-only transformer such as BERT reads an entire input sequence at once, attending in both directions, so every token's representation is informed by tokens both before and after it. It exists for tasks that require deep understanding of text, like classification, retrieval, and named entity recognition, rather than generating new text token by token.

BERT is trained with objectives like masked language modeling, where random words are hidden and the model learns to predict them from surrounding context, producing representations well suited to embeddings and search rather than open-ended generation.

How it worksThe mechanics

The full input sequence is fed in simultaneously and every token attends bidirectionally to every other token, producing a contextual vector per token; a pooled or per-token output then feeds a task-specific head such as a classifier, span extractor, or embedding projection. Because there is no causal mask, the model cannot generate text left-to-right, but it excels at producing dense representations for downstream tasks.

At a glanceSee it

Encoder (BERT) diagram
Encoder (BERT) diagram 1

How BERT actually learns — masked-language-model pretraining hides about 15 percent of tokens and predicts them from both sides, looping until the weights become a reusable encoder.

Encoder (BERT) diagram 2

The generate-versus-understand decision that places BERT's encoder-only design among the three transformer families.

When to use itWhere it fits

  • Text classification, sentiment analysis, and named entity recognition.
  • Producing dense embeddings for semantic search and retrieval pipelines.
  • Extractive question answering, where the answer is a span in the source text.
  • Fine-tuning a compact model for a narrow, well-defined understanding task.

When NOT to use itLimits & anti-patterns

  • Open-ended text generation or chat, where a decoder model is required instead.
  • Long free-form composition tasks like drafting emails or code, which encoders are not built for.
  • Zero-shot instruction following, where encoder-only models generally underperform instruction-tuned decoders.

Trade-offsAdvantages & costs

Advantages
  • Strong, efficient representations for classification and retrieval tasks.
  • Bidirectional context improves understanding versus left-to-right-only models.
  • Smaller and cheaper to run than large generative decoders for the same task.
  • Well suited to fine-tuning on modest labeled datasets.
Trade-offs & costs
  • Cannot generate free-form text at all.
  • Requires task-specific fine-tuning or heads rather than zero-shot prompting.
  • Less flexible for the multi-task, instruction-following use cases now common in production.
  • Overshadowed in mindshare by generative decoder models.

ExampleIn the real world

Google Search has used BERT-based models since 2019 to better understand the intent behind search queries before ranking results.

ToolsHow to implement it

  • Hugging Face Transformershosts BERT, RoBERTa, and DeBERTa checkpoints.
  • sentence-transformersencoder models fine-tuned for embedding and semantic search.
  • spaCyintegrates transformer encoders for NER and classification pipelines.
  • Cohere Embedproduction embedding API built on encoder-style architectures.

Cost & effortWhat it takes

Cheap to fine-tune and run compared to generative LLMs; low inference latency; needs labeled data for the target task but far less than training a model from scratch.

A living map of modern AI — kept current every morning