Home › The Road to LLMs › Classical ML
🛤️ · Foundations

Classical ML

Algorithms that learn statistical patterns from labeled data instead of hand-coded rules.

In one line

Classical ML replaced hand-written rules with algorithms that learn patterns directly from examples.

ConceptWhat it is

Classical machine learning covers algorithms like linear regression, decision trees, random forests, and support vector machines that learn statistical patterns from structured, labeled data instead of relying on hand-coded logic. It emerged in the 1990s and 2000s as data became more available and symbolic systems hit their ceiling.

It exists because most real-world problems have patterns too complex or too numerous for a human to enumerate as rules, but simple enough that a model can learn them from modest amounts of tabular data.

How it worksThe mechanics

A model is given labeled examples of inputs and outputs, then an optimization algorithm adjusts internal parameters, like tree splits or regression weights, to minimize prediction error, producing a function that generalizes to new, unseen inputs.

At a glanceSee it

Classical ML diagram
Classical ML diagram 1

Classical ML is a toolbox rather than one algorithm — the target type and the need for a readable rationale steer which family you reach for.

Classical ML diagram 2

Model complexity is a tug-of-war between fitting the training data and generalising, and the gap between training and validation error is the dial that tells you which way to turn.

When to use itWhere it fits

  • Structured tabular data with a clear target column, like churn or price prediction.
  • Limited data, on the order of thousands not millions of rows.
  • Interpretability matters, such as credit scoring or medical risk models.
  • Fast, cheap inference is needed on modest hardware.

When NOT to use itLimits & anti-patterns

  • Unstructured data like images, audio, or free text where feature engineering by hand is impractical.
  • Extremely large datasets where deep learning captures more signal.
  • Problems requiring the model to generate novel content rather than predict a label.

Trade-offsAdvantages & costs

Advantages
  • Interpretable, especially trees and linear models.
  • Fast to train and cheap to run.
  • Works well on small to medium datasets.
  • Mature tooling and strong theoretical understanding.
Trade-offs & costs
  • Requires manual feature engineering.
  • Struggles with unstructured, high-dimensional data.
  • Plateaus in accuracy compared to deep learning at scale.
  • Many algorithms assume independence or linearity that real data violates.

ExampleIn the real world

Capital One and most consumer banks still use gradient-boosted trees, like XGBoost, for credit-risk scoring because regulators require explainable decisions, not just accurate ones.

ToolsHow to implement it

  • scikit-learnthe standard Python library for classical algorithms and pipelines.
  • XGBoostindustry-standard gradient boosting for tabular competitions and production.
  • LightGBMfaster gradient boosting for large tabular datasets.
  • SHAPexplains individual predictions from tree-based models.

Cost & effortWhat it takes

Cheap to train, often minutes on a laptop CPU; inference is near-instant; the main cost is data engineering and feature preparation, not compute.

A living map of modern AI — kept current every morning