Classification answers which bucket something belongs in, like spam or not spam.
ConceptWhat it is
Classification is a supervised learning task where the model predicts a discrete category, such as spam versus not spam, or which of several disease types a patient likely has, rather than a continuous number.
It exists because a huge share of practical decisions are categorical: approve or deny, fraud or legitimate, this class or that one, and classification models give a probability-backed answer to exactly that kind of question.
How it worksThe mechanics
The model learns a decision boundary, a rule separating the feature space into regions, from labeled training examples, then for a new input it outputs either a predicted class directly or a probability for each possible class, and the highest-probability class is chosen.
At a glanceSee it
A taxonomy of classification by label structure — binary and multiclass each assign one label per item, while multilabel lets several tags apply at once.
The four outcomes of a binary prediction — true and false positives and negatives — showing why a miss and a false alarm rarely carry equal cost.
When to use itWhere it fits
- Deciding between a fixed set of discrete outcomes, like fraud detection.
- Content moderation, tagging text or images into predefined categories.
- Medical diagnosis support, predicting a condition from symptoms or scans.
- Any decision that ultimately reduces to picking one option from a known list.
When NOT to use itLimits & anti-patterns
- The target is a continuous number, like price or temperature, where regression is correct.
- Categories are not known in advance, where unsupervised clustering fits better.
- Class boundaries are extremely imbalanced or ill-defined without careful handling.
Trade-offsAdvantages & costs
Advantages
- Directly answers yes-or-no and multi-category decisions.
- Outputs calibrated probabilities useful for ranking or thresholds.
- Well-established metrics like precision, recall, and F1 for evaluation.
- Wide range of proven algorithms from logistic regression to deep nets.
Trade-offs & costs
- Struggles with severe class imbalance without special handling.
- A wrong category assignment can be costly in high-stakes settings.
- Requires labeled examples for every class of interest.
- Does not capture ordinal or continuous relationships between classes.
ExampleIn the real world
Gmail's spam filter is a classification model, scoring every incoming email as spam or not spam based on learned patterns from billions of labeled messages.
ToolsHow to implement it
- scikit-learnlogistic regression, random forests, and SVMs for classification.
- XGBoosttop performer for tabular classification tasks.
- PyTorchused for image and text classification with deep networks.
- Ragasevaluation metrics used to score classification-style outputs in ML pipelines.
Cost & effortWhat it takes
Training cost ranges from trivial, seconds for logistic regression, to substantial for deep classifiers on images; inference is typically fast and cheap once deployed.