📊 · Foundations

Regression

Predicting a continuous numeric value rather than a discrete category.

In one line

Regression predicts a number, like a price or a temperature, not a category.

ConceptWhat it is

Regression is a supervised learning task where the model predicts a continuous numeric value, such as a house price, a demand forecast, or a temperature, rather than a discrete class.

It exists because many real business questions are fundamentally about how much or how many, not which category, and regression models give a calibrated numeric estimate along with a sense of expected error.

How it worksThe mechanics

The model learns a function mapping input features to a numeric output by minimizing the difference, often squared error, between its predictions and the true numeric values in the training data, then applies that learned function to new inputs.

At a glanceSee it

Regression diagram
Regression diagram 1

How fitting actually works — predict, measure the error, nudge the parameters, and repeat until the error stops falling.

Regression diagram 2

Choosing the scorecard — squared error to punish big misses, absolute error for robustness, and R-squared for the share of variance explained.

When to use itWhere it fits

  • Forecasting a continuous quantity, like sales, demand, or price.
  • Estimating risk scores or probabilities on a continuous scale.
  • Any scenario where how much matters more than which category.
  • Time-series style numeric forecasting with tabular features.

When NOT to use itLimits & anti-patterns

  • The output is inherently categorical, where classification is the right tool.
  • The relationship between inputs and output is highly non-numeric or symbolic.
  • Data has extreme outliers that distort standard error-minimizing regression without robust methods.

Trade-offsAdvantages & costs

Advantages
  • Gives a precise, continuous estimate rather than a coarse bucket.
  • Well-understood statistical foundations and diagnostics.
  • Easy to measure error with metrics like RMSE or MAE.
  • Works well with both simple linear models and complex ensembles.
Trade-offs & costs
  • Sensitive to outliers unless robust methods are used.
  • Assumes a reasonably smooth relationship between inputs and output.
  • Extrapolation beyond the training data range is unreliable.
  • Choosing the right error metric matters and is often overlooked.

ExampleIn the real world

Zillow's Zestimate uses regression models trained on property features and sales history to predict a specific dollar value for millions of homes.

ToolsHow to implement it

  • scikit-learnlinear regression, ridge, lasso, and ensemble regressors.
  • XGBooststrong performer for tabular regression tasks.
  • statsmodelsdetailed statistical diagnostics for regression models.
  • Prophetpopular library for time-series regression forecasting.

Cost & effortWhat it takes

Training is typically cheap, from milliseconds for linear models to minutes for boosted ensembles; inference is fast; main effort goes into feature engineering and validation.

A living map of modern AI — kept current every morning