A foundation model is one big pre-trained model that many different products build on.
DefinitionWhat it means
A foundation model is a large model, such as GPT, Claude, or Gemini, pre-trained on broad data at massive scale and designed to serve as a base for many downstream applications through fine-tuning, prompting, or retrieval, rather than being built for a single narrow task. The term, coined by Stanford researchers, emphasizes that these models function as general-purpose infrastructure rather than task-specific tools.
Why it mattersWhy you should care
Foundation models reshaped the AI industry's economics: instead of every company training its own model per task, most now build on top of a small number of foundation models via APIs or fine-tuning, concentrating enormous training cost among a few labs while distributing value creation broadly. For buyers, choosing a foundation model, and understanding its licensing, is now one of the first strategic decisions in any AI product.
At a glanceSee it
The word pre-trained hides a multi-stage build — self-supervised pretraining yields a raw base model, then a reward-model loop aligns it, and alignment only masks rather than erases the base model's behavior.
Adapting a foundation model is a decision, not a default — retrieval adds facts, fine-tuning adds behavior, prompting just steers, and reaching for fine-tuning to inject facts freezes stale knowledge in place.
Where you see itIn the wild
- Vendor comparisons between GPT, Claude, and Gemini as foundation models
- Procurement discussions about which foundation model to standardize on
- Papers describing foundation model as a category distinct from narrow AI