Overfitting is a model doing great on practice data and poorly on the real thing.
DefinitionWhat it means
Overfitting occurs when a model learns the noise and idiosyncrasies of its training data so closely that it fails to generalize to new, unseen data. It typically shows up as a large gap between training accuracy and validation or test accuracy, and it is more likely with complex models trained on too little or too repetitive data. Common defenses include regularization, dropout, early stopping, and holding out a validation set.
Why it mattersWhy you should care
Overfitting is the reason a model can look excellent in a demo built on curated examples and then fail badly in production against real, messy user input. Product and data teams guard against it with rigorous train and test splits, cross-validation, and monitoring live performance against offline benchmarks, because an overfit model creates false confidence right before launch.
At a glanceSee it
Detecting overfitting from the widening gap between train and validation error, then looping through remedies before you ship.
The two families of fixes — constrain the model or add more data — that pull an overfit model back toward generalization.
Where you see itIn the wild
- A validation loss curve that diverges upward from training loss
- Model reviews probing how overfitting was detected and prevented
- Regularization and dropout settings in a training config