Home › Models
🤖 Models

LLMs & Foundation Models

Large models pre-trained on vast data, adaptable to almost any task.

At a glanceLLMs & Foundation Models

LLMs & Foundation Models diagram

Large models pre-trained on vast data, adaptable to almost any task.

LandscapeTypes & approaches

Click a highlighted type to open its own page — concept, use case, and diagram.

The one factEverything follows from next-token prediction

A large language model is a function that reads a sequence of tokens and produces a probability distribution over the next one. That is the whole mechanism. It is trained by being shown enormous quantities of text with the next token hidden, and adjusted until its guesses improve. Then it is run in a loop: predict a token, append it, predict again.

Almost every surprising behaviour these systems exhibit is a direct consequence of that sentence, and holding onto it will save you from most of the bad intuitions in circulation.

Why this matters practically

A model does not "look something up" and then phrase it. It produces the most plausible continuation. When the plausible continuation happens to be true, we call it knowledge; when it does not, we call it hallucination. These are the same operation with different luck. That is why you cannot fix hallucination with a stern instruction, and why grounding the model in retrieved text works so much better than asking it to be careful.

Four consequences you will meet in week one

It is fluent before it is correct. Fluency is what the objective directly optimises. Correctness is a side effect of the fact that true text is common in the training data. So a wrong answer arrives in the same confident register as a right one, and the prose style carries no signal about reliability. Human writing usually leaks doubt; this does not.

It is sensitive to how you ask. The distribution over the next token depends on every token before it, so the framing, the order of your examples and even your formatting move the output. This is not the model being temperamental. It is the model doing exactly the one thing it does, conditioned on a different input.

It has no memory between calls. Each request is scored fresh against whatever text you supply. Conversations feel continuous because the transcript is resent every turn. Everything the model "remembers" is something your application put back in front of it, which is why context management is an engineering problem and not a model feature.

Its knowledge has an edge. Pre-training ends on a date. Past that edge the model does not know that it does not know, because unfamiliarity is not represented anywhere in a next-token distribution. It will produce a fluent guess about last month with the same confidence it brings to a decade-old fact.

TrainingHow a base model becomes something you can talk to

The figure above is the pipeline, and the reason a raw base model and a helpful chat model behave so differently despite sharing most of their weights. Three stages, each buying something the previous one could not.

Pre-training buys capability

The model is trained on a very large corpus to predict the next token. This is where nearly all of the compute goes and where the capability comes from — grammar, facts, reasoning patterns, translation, code, the lot. It is also the stage you will never do. Pre-training a frontier model is a capital project, not an engineering task, and the output of this stage is what you are actually renting or downloading when you pick a model.

The size of a model and the amount of data it saw are not independent choices. The scaling literature established that loss falls predictably with compute, and the compute-optimal work that followed showed most large models of that era were substantially under-trained for their size — they should have seen far more data. The practical residue for you is that parameter count alone is a poor proxy for capability. A smaller, longer-trained model routinely beats a larger, under-trained one, which is why "how many billion parameters" is close to a meaningless question when comparing models from different labs and years.

Supervised fine-tuning buys obedience

A base model completes text; it does not answer questions. Ask one "What is the capital of France?" and a plausible continuation is another exam question, because that is what follows a question in a lot of scraped text. Supervised fine-tuning shows the model a much smaller set of curated instruction-and-response pairs until responding becomes the plausible continuation of being asked. Nothing new is learned about the world here. A behaviour is selected.

Preference tuning buys judgement

Obedient is not the same as helpful. The last stage optimises against human preference: show people several candidate responses, learn which they prefer, push the model toward those. The original approach trained a separate reward model and optimised against it with reinforcement learning; direct preference optimisation later showed you can hit a closely related objective by training on preference pairs directly, without standing up a reward model or an RL loop, which is why so much open-weight work uses it.

This stage is where tone, refusal behaviour, formatting habits and hedging come from. It is also the stage most responsible for the personality differences between models that are otherwise architecturally similar, and for the fact that a model can become less good at something it could previously do — preference tuning trades breadth for agreeableness, and that trade is not free.

The practical read

When a model refuses something reasonable, or hedges endlessly, or writes in an over-formatted register you did not ask for, you are meeting the preference-tuning stage, not a limitation of the underlying capability. Prompting can often recover it. When a model simply cannot do the task at all, you are meeting the pre-training stage, and no prompt will help.

The forkThe decision that shapes every decision after it

Open-weight versus closed API is the architecture-and-feasibility decision on this page. Nearly everything else — hosting, cost shape, data residency, how fast you can adopt a better model, whether fine-tuning is even available — is downstream of it. It deserves to be made deliberately and early, and it is reversible only at real expense.

 Closed APIOpen-weight
What you actually getAn endpoint and a rate limitA file of weights and a hosting problem
Cost shapePer token, scales with traffic, near-zero at idlePer GPU-hour, scales with capacity, paid while idle
Where your data goesTo someone else, under their termsWherever you put it
Adopting a better modelChange a string, usuallyRe-host, re-tune, re-evaluate
Worst failure modeDeprecation and silent behaviour driftYou now run a GPU fleet
Fine-tuningOnly if offered, only how it is offeredFully yours
Four model options positioned against cost to run and how much control you hold.

The fork is a position, not a sequence — so it is drawn as one. Placement is the shape of the trade-off, not a measurement of any specific model.

The honest case for closed

You get the strongest available capability with no capacity planning, no GPU procurement and no inference engineering, and you get it behind an interface stable enough that switching models is often a configuration change. At low and spiky volume the economics are not close: you pay nothing when nobody is using the product. For most teams building most features, this is the correct default, and choosing it is not a failure of ambition.

The honest case for open

Three reasons hold up. Data residency — some data legally or contractually cannot leave your boundary, and no amount of vendor assurance changes that. Volume — per-token pricing is linear while a GPU is fixed, so above some sustained throughput the fixed cost wins, and that crossover is a real number you can compute rather than argue about. Control — nobody deprecates your model, changes its behaviour underneath you, or rate-limits you at the worst moment.

Reasons that do not hold up as well as people expect: cost at low volume (an idle GPU is pure loss), and capability parity (the open frontier moves fast and is genuinely excellent, but a specific open model is not automatically a drop-in for a specific closed one on your task, and only your evaluation can tell you).

Do the arithmetic before the argument

The crossover between per-token and per-hour is computable from three numbers you already have: your sustained requests per second, your average tokens per request, and the hourly cost of hardware that serves your latency target. Teams routinely have this argument on principle for weeks when an hour of arithmetic would settle it. The costing pages carry the token maths; the frontier board carries the current per-token rates.

SpecificationsReading a model's spec sheet without being fooled

Every model ships with a short table of numbers. Each one is useful and each one is routinely over-read. Current values for the models worth considering live on the frontier board, dated and sourced; what follows is how to interpret them.

The numberWhat it tells youWhat it does not tell you
Context window The ceiling on prompt and response together. That quality holds up to it. Attention over a long context is not uniform — material mid-prompt is attended to less reliably than at either end.
Maximum output The real limit on generated length. Usually far smaller than the context window. Nothing about input capacity. This, not the window, is what truncates long structured documents.
Input price What you pay on every call, for the whole prompt. The real bill, until you multiply by how often prompts grow. A context window is a budget spent per request, not a capability owned.
Output price Typically several times the input rate. — and it inverts a common intuition: a long prompt with a short answer is often cheaper than a short prompt with a long one.
Cache-hit input price What a large stable prefix costs on repeat — a system prompt, a fixed document, a tool schema. Whether your workload actually has a stable prefix. Where it does, this is often the difference between affordable and not.
Knowledge cutoff Where pre-training ended. Much that matters. Anything whose answer changes should come from retrieved context regardless. If correctness depends on this date, the architecture is wrong, not the date.
Prices belong in one place

This page deliberately quotes none. Per-token rates change often enough that a number written into prose is a future lie, and a second copy of a price is a copy that will eventually disagree with the first. The frontier board carries the current figures with the date each was verified and the source it came from. When those two would ever conflict, there is only one of them, so they cannot.

SelectionThe four shapes you are actually choosing between

Vendor naming obscures more than it reveals, and it changes every few months. Underneath the branding there are four shapes, and knowing which one you need eliminates most of the catalogue immediately.

ShapeUse it whenThe trade you are making
FrontierThe task is genuinely hard, or you do not yet know how hard it isHighest cost and latency; the safest place to discover whether a thing is possible at all
Mid-tierThe task is understood and quality is provenWhere most production traffic should end up; materially cheaper for a small quality step down
SmallNarrow, high-volume, well-specified work — classify, extract, route, rewriteWill not generalise past its task; often beats a frontier model there anyway, especially fine-tuned
ReasoningMulti-step problems where a wrong answer is expensiveSpends tokens and seconds thinking before answering; wasteful on anything simple

On reasoning models specifically

These are trained to produce an extended internal working-out before committing to an answer, building on the finding that prompting a model to reason step by step improves accuracy on multi-step problems. The gain is real on maths, code, planning and anything with a chain of dependent steps. It is close to zero on retrieval, summarisation, formatting and classification — and on those tasks you are paying for thinking tokens and waiting for them.

The mistake to avoid is routing everything to one because it benchmarks well. Reasoning is a tool for a class of problem, not a general upgrade, and applying it uniformly is one of the easiest ways to multiply your bill and your latency without moving your quality.

On small models specifically

The most reliably under-used option. When the task is narrow and the volume is high, a small model — particularly one fine-tuned on a few thousand of your own examples — will often match or beat a frontier model on that task at a fraction of the cost and latency, and can run somewhere a frontier model cannot. The catch is that it will not generalise past what you trained it for, so this only works when the task is genuinely stable.

Ground levelWhat you actually build

One pipeline in, many models out. Open or closed, big or small is a trade-off you keep re-making, not a decision you land once.

One pipeline in, many models out. Open or closed, big or small is a trade-off you keep re-making, not a decision you land once.

ArchitectureThe routing seam, or how not to marry a model

The frontier moves monthly. Any model you choose today is a temporary answer, and the most valuable thing you can build is not a good choice but a cheap way to change your mind.

Concretely, that means no part of your application should name a model except one. Calls go through a seam that owns model selection, and everything above it asks for a capability rather than a product.

A request passing through the routing seam, tried on a cheap model first and escalated to a frontier model only when needed.

Who is asked what, in what order. The application never learns which model answered.

Once that seam exists you get several things at once that are painful to retrofit later:

Routing by difficulty. Most traffic in most products is easy. Sending it all to your strongest model is the most common and most expensive mistake in the field. A classifier or a simple heuristic in front of the seam, sending easy work to a cheap model and escalating the rest, is frequently the single largest cost lever available, and it is invisible to your users when it works.

Evaluating a swap. With a seam you can run a new model against your recorded traffic and your eval set before committing. Without one, "should we switch?" is a matter of opinion, and switching is a rewrite.

Surviving deprecation. Closed models are retired on the vendor's schedule. With a seam this is a configuration change made calmly. Without one it is an incident.

Falling back. Providers have outages and rate limits. The seam is the only sane place to degrade to a second choice rather than to an error page.

The failure this prevents

The expensive version of this mistake is not the model name appearing in fifty files. It is the prompts, the output parsing and the retry logic having quietly grown around one model's particular habits — its formatting tendencies, its refusal patterns, how it behaves when asked for JSON. That coupling is invisible until you try to swap, at which point the swap is not a configuration change and never was.

MethodHow to choose, in order

The order matters more than the individual steps, because each one cheaply eliminates options that would otherwise waste the next one's effort.

StepWhat you doWhy it sits here and not later
1 · Write the quality bar State what a good answer is and how you will recognise one, before looking at any model. Skip it and every comparison afterwards is decided on impression. This is why so many model bake-offs end in a shrug.
2 · Eliminate on hard constraints Data residency, latency ceiling, on-device, procurement. These are binary and remove most of the field in one pass, so you never benchmark a model you were never allowed to use.
3 · Prototype on something strong Use a frontier model to find out whether the task is achievable at all. Answering "is this possible?" and "what should this cost?" at once is how teams conclude a feature is impossible when it was only under-resourced.
4 · Only then optimise down Walk the ladder — mid-tier, small, fine-tuned small — and stop at the cheapest model that still clears the bar. It needs the eval set from step 1 to be a measurement rather than a guess. This is the step that pays for the other four.
5 · Re-run on a schedule Put the current field through the eval set you already have. The answer expires. A model released after your evaluation may be cheaper and better, and the seam makes finding out nearly free.

FailureWhere this goes wrong

The mistakeWhat it looks like from insideWhat it actually costs
Choosing on benchmarks A leaderboard column decides the shortlist and then the winner. They measure something, but not your task, prompts or data. A small eval set of your own real examples outranks every public number.
Choosing on vibes after ten prompts "We tried both and this one felt better." Ten prompts is a demo. Variance from how you phrased them exceeds the variance between the models, so it measures your prompting.
Parameter count as capability Ranking models by billions. Training volume, tuning quality and architecture move capability as much as size. Across labs and years the comparison is close to meaningless.
Frontier prices for easy traffic One model configured for everything. Usually the largest unexamined line in an AI budget, and invisible because nothing is broken.
Assuming a swap is free "It is the same API, we can switch any time." The surface may be identical while behaviour is not. A provider-compatible endpoint does not make two models interchangeable.
Fine-tuning to fix knowledge "The model does not know our products, so we will train it." Fine-tuning adjusts behaviour, not facts. If it must know something specific and current, retrieve it.

ReferencesSources

The mechanisms described above, in the primary literature. Volatile figures — prices, context windows, benchmark scores — are deliberately absent from this page and live on the frontier board instead, each carrying the date it was verified and the source it came from.

  1. Vaswani et al. — Attention Is All You Need (the transformer)
  2. Kaplan et al. — Scaling Laws for Neural Language Models
  3. Hoffmann et al. — Training Compute-Optimal Large Language Models (Chinchilla)
  4. Ouyang et al. — Training language models to follow instructions with human feedback
  5. Rafailov et al. — Direct Preference Optimization
  6. Wei et al. — Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
  7. Hu et al. — LoRA: Low-Rank Adaptation of Large Language Models

GlossaryKey terms

CheckCheck your understanding

Why does a model hallucinate confidently rather than expressing doubt?

Because it produces the most plausible continuation, and plausibility is what the training objective rewards. Truth is a frequent side effect of plausible text, not a separate thing the model tracks - so nothing in the mechanism distinguishes a correct continuation from an incorrect one. Confidence is a property of the prose, not a signal about reliability.

A base model and a chat model share most of their weights. Why do they behave so differently?

Because supervised fine-tuning and preference tuning select a behaviour from capability that pre-training already established. The knowledge is largely the same; what changed is which continuation is plausible when the input looks like a question.

When does open-weight actually beat a closed API?

When data cannot leave your boundary, when sustained volume pushes you past the crossover between per-token and per-hour cost, or when you need control over deprecation and behaviour drift. Low-volume cost is not on that list - an idle GPU is pure loss.

Your bill is too high. Where do you look first?

At routing, not at the model list. Most traffic is easy and is probably going to your strongest model. After that look at input tokens, since you pay them on every call and prompts grow quietly.

Why is 'how many parameters?' a weak question?

Because a smaller model trained on far more data routinely outperforms a larger under-trained one, and tuning quality moves capability independently of size. Parameter count is one input to capability, not a measure of it.

What changedWhat changed here

MeasuredCounted in the daily brief, with every item listed
Written inYou approved this and it changed the page
  • Updated this page Google now offers speech-to-speech models alongside OpenAI's, giving voice-interface builders a second vendor to price and benchmark against.

    Google DeepMind · 15 Sep 2026 · source

  • Updated this page Research using expert masking shows that MoE layers differ in importance, which informs understanding of mixture-of-experts models.

    arXiv cs.AI · 16 Aug 2026 · source

  • Updated this page Research suggests LLMs develop a modular cognitive architecture analogous to brain specialization.

    arXiv cs.AI · 16 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning