MACHINE LEARNING / CLUSTERING

HIERARCHICAL CLUSTERING

Build a family tree of data by repeatedly merging the closest clusters.

UnsupervisedClusteringHierarchical
Saved only in this browser

01 Overview

02 The Problem

When the number of groups is uncertain, hierarchical clustering builds a nested history of merges so the data can be inspected at several granularities.

03 Why It Matters

The dendrogram preserves a record of which clusters joined and at what distance. Cutting it at different heights gives different numbers of groups.

04 Intuition

Start with every point alone. Repeatedly merge the closest pair of clusters until one family tree remains.

05 Mathematical Foundation

A linkage rule defines distance between clusters. Single linkage uses the closest pair across two clusters, while other rules use averages or furthest pairs.

06 The Equation

\[d_{single}(A,B)=\min_{a\in A,b\in B}d(a,b)\]
  • \(A,B\) candidate clusters
  • \(d(a,b)\) point distance
  • \(d_{single}\) linkage distance

07 How It Learns

  1. Initialise one cluster per point.
  2. Find the closest pair under the linkage rule.
  3. Merge that pair.
  4. Record the distance and repeat.

08 Algorithm

Singleton clusters
↓
Closest merge
↓
Dendrogram

09 Visual Explanation

A quick visual summary of how this model sees data and makes its prediction.

10 Worked Example

A dendrogram can be cut just before a large jump in merge distance, treating that jump as evidence that distinct groups are being joined.

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning