Home › Data Science & ML Fundamentals › Key term › Overfitting
Key term · Foundations

Overfitting

When a model memorizes training data instead of learning general patterns.

In one line

Overfitting is a model doing great on practice data and poorly on the real thing.

DefinitionWhat it means

Overfitting occurs when a model learns the noise and idiosyncrasies of its training data so closely that it fails to generalize to new, unseen data. It typically shows up as a large gap between training accuracy and validation or test accuracy, and it is more likely with complex models trained on too little or too repetitive data. Common defenses include regularization, dropout, early stopping, and holding out a validation set.

Why it mattersWhy you should care

Overfitting is the reason a model can look excellent in a demo built on curated examples and then fail badly in production against real, messy user input. Product and data teams guard against it with rigorous train and test splits, cross-validation, and monitoring live performance against offline benchmarks, because an overfit model creates false confidence right before launch.

At a glanceSee it

Overfitting diagram
Overfitting diagram 1

Detecting overfitting from the widening gap between train and validation error, then looping through remedies before you ship.

Overfitting diagram 2

The two families of fixes — constrain the model or add more data — that pull an overfit model back toward generalization.

Where you see itIn the wild

  • A validation loss curve that diverges upward from training loss
  • Model reviews probing how overfitting was detected and prevented
  • Regularization and dropout settings in a training config
A living map of modern AI — kept current every morning