Home › Foundations
🛤️ Foundations

The Road to LLMs

How seven decades of AI led to models that understand and generate language.

OverviewWhat it is

LLMs did not appear overnight. Every era of AI removed a limitation of the one before it. Early symbolic systems only knew what a human explicitly coded. Machine learning let systems learn patterns from data instead of rules. Deep learning let networks learn the features themselves. The Transformer (2017) let a model read a whole sequence in context; pre-training let one model serve many tasks; scale unlocked surprising new abilities; and alignment (RLHF) finally made models follow human intent.

At a glanceThe Road to LLMs

The Road to LLMs diagram

Left to right: each era removed a limitation of the one before it, until models learned language itself.

MechanicsHow it works

The through-line is a steady trade of hand-coding for learning. Rules gave way to statistical learning; hand-crafted features gave way to learned representations; task-specific models gave way to general foundation models adapted cheaply. The 2017 Transformer is the hinge: self-attention made both scale and quality possible at once.

Ground levelWhat you actually build

Two eras and one hinge between them. Knowing which era's tool fits the problem in front of you is a design decision, not trivia.

Two eras and one hinge between them. Knowing which era's tool fits the problem in front of you is a design decision, not trivia.

LandscapeTypes & approaches

Click a highlighted type to open its own page — concept, use case, and diagram.

FeasibilityArchitecture & feasibility

Architecture & feasibility

  • Every era still ships: a fraud model (classical ML), a vision CNN, and an LLM can co-exist in one product. Architecture feasibility means choosing the right tool per sub-problem, not defaulting to the newest.
  • The Transformer's parallelism is why LLMs are GPU-bound and why context length drives cost — a fact that shapes every downstream design decision.

In practiceWhat it means for building

Use this arc to set expectations: LLMs are powerful pattern machines, not reasoning oracles. Knowing the lineage helps you separate genuine capability from hype.

Treat 'AI' as a portfolio of techniques. The architecture question is which era's tool fits each sub-problem, and how they compose into one reliable system.

GlossaryKey terms

CheckCheck your understanding

Walk me through how we got to LLMs.

Rules to classical ML to deep learning to embeddings to the Transformer (2017) to pre-training to scale (GPT-3) to alignment (ChatGPT) to multimodal and agents. Each step traded hand-coding for learning.

What changed in 2017?

The Transformer replaced sequential RNNs with self-attention, which reads the whole sequence in parallel and captures long-range context, enabling both scale and quality.

Is deep learning always the right choice?

No. For structured, well-defined problems a classical model is cheaper, faster, and more explainable. The lineage helps you pick deliberately.

What changedWhat changed here

Nothing in the daily brief has touched this page since 2026-09-25. The sweep runs every morning and checks every page on this site; when it finds something for this one, it lands here.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning