📊 · Foundations

Clustering

Grouping similar data points together without any predefined category labels.

In one line

Clustering automatically groups similar items together without anyone telling it what the groups should be.

ConceptWhat it is

Clustering is an unsupervised technique that groups data points so that items within a group are more similar to each other than to items in other groups, without any predefined category labels to guide it.

It exists because many datasets have natural structure, customer segments, gene expression patterns, document topics, that is not known in advance, and clustering surfaces that structure directly from the data's own geometry.

How it worksThe mechanics

The algorithm measures similarity or distance between data points in feature space, then iteratively assigns points to groups, or builds a hierarchy of groups, so that overall within-group similarity is maximized and between-group similarity is minimized.

At a glanceSee it

Clustering diagram
Clustering diagram 1

Which clustering family to run is itself a decision driven by what you already know about the data—how many groups exist and whether their shape and density vary.

Clustering diagram 2

With no labels to check against, the number of clusters is chosen internally—sweep k, keep the value whose silhouette is tightest and most separated, and accept that a flat curve means the data may not cluster at all.

When to use itWhere it fits

  • Customer or user segmentation without predefined categories.
  • Exploratory analysis of a new dataset to find natural groupings.
  • Anomaly detection, treating small or distant clusters as outliers.
  • Document or topic grouping when categories are not known beforehand.

When NOT to use itLimits & anti-patterns

  • A known, fixed set of target categories exists, where classification is more direct.
  • Clusters need to be objectively validated against ground truth, which clustering cannot provide.
  • Data has no real underlying group structure, which produces meaningless clusters.

Trade-offsAdvantages & costs

Advantages
  • Requires no labeled data at all.
  • Surfaces structure a human analyst might miss.
  • Useful as a first exploratory step on any new dataset.
  • Many algorithms scale to large datasets efficiently.
Trade-offs & costs
  • No single correct answer, results depend on algorithm and parameter choices.
  • Choosing the right number of clusters is often subjective.
  • Hard to evaluate objectively without domain expert review.
  • Sensitive to feature scaling and irrelevant dimensions.

ExampleIn the real world

Retailers like Sephora use clustering on purchase history to segment customers into groups such as value shoppers or luxury loyalists for targeted marketing.

ToolsHow to implement it

  • scikit-learnk-means, DBSCAN, and hierarchical clustering implementations.
  • HDBSCANrobust density-based clustering for irregular shapes.
  • UMAPdimensionality reduction often paired with clustering for visualization.
  • Faissefficient similarity search useful for clustering at large scale.

Cost & effortWhat it takes

Generally cheap and fast on datasets up to millions of rows; cost rises with high dimensionality or very large scale; main effort is in choosing features and validating cluster quality.

A living map of modern AI — kept current every morning