Home › The Road to LLMs › Scaling laws
🛤️ · Foundations

Scaling laws

Predictable relationships showing how model performance improves with more data, parameters, and compute.

In one line

Scaling laws show that bigger models trained on more data and compute reliably get better, in predictable, measurable amounts.

ConceptWhat it is

Scaling laws are empirical relationships, first rigorously documented by OpenAI and later refined by DeepMind's Chinchilla paper, showing that a model's loss decreases in a smooth, predictable curve as you increase parameters, training data, and compute together.

They exist because they let labs forecast how much a bigger model will improve before spending millions of dollars training it, and they revealed that many early models were undertrained relative to their size, reshaping how compute budgets are allocated.

How it worksThe mechanics

Researchers train a series of smaller models at different sizes and data volumes, plot loss against compute on a log-log scale, and fit a power-law curve that extrapolates how much a much larger training run should improve performance before it is actually run.

At a glanceSee it

Scaling laws diagram
Scaling laws diagram 1

The Chinchilla decision — for a fixed compute budget, balancing parameters against training tokens beats pouring it all into raw size, though inference economics can still push you to overtrain a smaller model.

Scaling laws diagram 2

Anatomy of the loss curve — test loss is a shrinking power-law term stacked on a fixed entropy floor, which is why returns diminish near the floor and why a smooth curve can still miss capabilities that emerge in jumps.

When to use itWhere it fits

  • Planning a large training run and needing to forecast expected performance gains.
  • Deciding the optimal ratio of model size to training data, as in Chinchilla-style scaling.
  • Justifying compute budgets to leadership before an expensive training run.
  • Comparing whether more data or more parameters will help a given model more.

When NOT to use itLimits & anti-patterns

  • Small fine-tuning jobs where scaling laws barely apply and empirical testing is faster.
  • Tasks bottlenecked by data quality or evaluation design rather than raw scale.
  • Situations where scaling has plateaued for the task and architecture or data changes matter more.

Trade-offsAdvantages & costs

Advantages
  • Gives a data-driven forecast before committing large compute budgets.
  • Revealed that data volume matters as much as parameter count.
  • Helps avoid wasteful undertrained or overtrained models.
  • Widely validated across many model families and modalities.
Trade-offs & costs
  • Curves can break down at extreme scales or for new architectures.
  • Does not account for data quality, only quantity.
  • Requires expensive small-scale experiments to fit the curve accurately.
  • Says nothing about emergent capabilities appearing at specific thresholds.

ExampleIn the real world

DeepMind's Chinchilla paper showed that Gopher, a 280-billion-parameter model, was undertrained, and a 70-billion-parameter model trained on more data, Chinchilla, outperformed it, reshaping how labs budget compute.

ToolsHow to implement it

  • Weights and Biasestracks loss curves across scaling experiments.
  • Rayorchestrates the distributed training runs scaling studies require.
  • Chinchilla methodologythe reference scaling framework most labs now follow.
  • NVIDIA Megatron toolinginfrastructure for running scale experiments.

Cost & effortWhat it takes

Fitting scaling laws itself requires running many mid-sized training jobs, a meaningful compute investment, but it prevents far costlier mistakes on the full-scale run.

A living map of modern AI — kept current every morning