Regression predicts a number, like a price or a temperature, not a category.
ConceptWhat it is
Regression is a supervised learning task where the model predicts a continuous numeric value, such as a house price, a demand forecast, or a temperature, rather than a discrete class.
It exists because many real business questions are fundamentally about how much or how many, not which category, and regression models give a calibrated numeric estimate along with a sense of expected error.
How it worksThe mechanics
The model learns a function mapping input features to a numeric output by minimizing the difference, often squared error, between its predictions and the true numeric values in the training data, then applies that learned function to new inputs.
At a glanceSee it
How fitting actually works — predict, measure the error, nudge the parameters, and repeat until the error stops falling.
Choosing the scorecard — squared error to punish big misses, absolute error for robustness, and R-squared for the share of variance explained.
When to use itWhere it fits
- Forecasting a continuous quantity, like sales, demand, or price.
- Estimating risk scores or probabilities on a continuous scale.
- Any scenario where how much matters more than which category.
- Time-series style numeric forecasting with tabular features.
When NOT to use itLimits & anti-patterns
- The output is inherently categorical, where classification is the right tool.
- The relationship between inputs and output is highly non-numeric or symbolic.
- Data has extreme outliers that distort standard error-minimizing regression without robust methods.
Trade-offsAdvantages & costs
Advantages
- Gives a precise, continuous estimate rather than a coarse bucket.
- Well-understood statistical foundations and diagnostics.
- Easy to measure error with metrics like RMSE or MAE.
- Works well with both simple linear models and complex ensembles.
Trade-offs & costs
- Sensitive to outliers unless robust methods are used.
- Assumes a reasonably smooth relationship between inputs and output.
- Extrapolation beyond the training data range is unreliable.
- Choosing the right error metric matters and is often overlooked.
ExampleIn the real world
Zillow's Zestimate uses regression models trained on property features and sales history to predict a specific dollar value for millions of homes.
ToolsHow to implement it
- scikit-learnlinear regression, ridge, lasso, and ensemble regressors.
- XGBooststrong performer for tabular regression tasks.
- statsmodelsdetailed statistical diagnostics for regression models.
- Prophetpopular library for time-series regression forecasting.
Cost & effortWhat it takes
Training is typically cheap, from milliseconds for linear models to minutes for boosted ensembles; inference is fast; main effort goes into feature engineering and validation.