HIERARCHICAL CLUSTERING
Build a family tree of data by repeatedly merging the closest clusters.
01 Overview
02 The Problem
When the number of groups is uncertain, hierarchical clustering builds a nested history of merges so the data can be inspected at several granularities.
03 Why It Matters
The dendrogram preserves a record of which clusters joined and at what distance. Cutting it at different heights gives different numbers of groups.
04 Intuition
Start with every point alone. Repeatedly merge the closest pair of clusters until one family tree remains.
05 Mathematical Foundation
A linkage rule defines distance between clusters. Single linkage uses the closest pair across two clusters, while other rules use averages or furthest pairs.
06 The Equation
- \(A,B\) candidate clusters
- \(d(a,b)\) point distance
- \(d_{single}\) linkage distance
07 How It Learns
- Initialise one cluster per point.
- Find the closest pair under the linkage rule.
- Merge that pair.
- Record the distance and repeat.
08 Algorithm
09 Visual Explanation
A quick visual summary of how this model sees data and makes its prediction.
10 Worked Example
A dendrogram can be cut just before a large jump in merge distance, treating that jump as evidence that distinct groups are being joined.