Scaling laws say bigger models trained on more data get reliably better.
DefinitionWhat it means
Scaling laws are empirical relationships, mapped in detail by researchers at OpenAI and DeepMind, showing that a model's loss decreases in a smooth, predictable way as training data volume, parameter count, and compute budget all increase together. These curves let researchers extrapolate performance before running an expensive training job, which is why frontier labs plan multi-hundred-million dollar runs with confidence rather than guesswork.
Why it mattersWhy you should care
Scaling laws turned model-building from trial and error into an engineering discipline: teams can forecast whether a 10x compute increase justifies the cost before spending it. For buyers evaluating AI vendors, scaling laws explain why newer, larger models keep beating older ones on benchmarks, and why compute budgets, not just clever algorithms, remain a primary lever for capability.
At a glanceSee it
A fixed compute budget forces a model-size versus data split — Chinchilla’s balanced point of about twenty tokens per parameter is loss-optimal, a giant undertrained model wastes the budget, and a data wall can put that optimum out of reach.
Opening the loss curve reveals total loss as an irreducible entropy floor plus a reducible term that falls as a power law — which is why loss versus compute traces a straight log-log line that eventually flattens and never reaches zero.
Where you see itIn the wild
- Chinchilla and GPT scaling-law papers cited in model release notes
- Capacity planning discussions before a large training run
- Investor decks justifying compute spend for a next-generation model