MACHINE LEARNING / CLUSTERING

K-MEANS CLUSTERING

Group unlabeled points around K moving centroids until assignments stop changing.

UnsupervisedClusteringCentroid-based
Saved only in this browser

01 Overview

02 The Problem

Suppose customer records have no labels. K-Means groups nearby records into a chosen number of segments by repeatedly assigning points to their nearest centre and moving each centre to its group mean.

03 Why It Matters

Clustering turns an unlabelled dataset into a first map of its structure. It is useful for segmentation, compression, anomaly review, and exploratory analysis.

04 Intuition

Place k movable centres on the map. Every point joins its closest centre, then every centre moves to the average of its members. Alternating those two steps settles into a locally good partition.

05 Mathematical Foundation

The objective is within-cluster squared distance. Euclidean distance makes the mean the centre that minimises the total squared error for its assigned points.

06 The Equation

\[ J = \sum_{k=1}^{K}\sum_{x_i\in C_k}\|x_i-\mu_k\|^2 \]
  • \(C_k\) points assigned to cluster k
  • \(\mu_k\) mean of cluster k
  • \(J\) total within-cluster inertia

07 How It Learns

  1. Choose k initial centres.
  2. Assign each point to its nearest centre.
  3. Replace each centre with its assigned mean.
  4. Repeat until assignments stop changing.

08 Algorithm

Initial centres
↓
Assign
↓
Update means
↓
Converged clusters

09 Visual Explanation

A quick visual summary of how this model sees data and makes its prediction.

10 Worked Example

With k = 2, points near (1, 1) and (8, 8) quickly pull their centres toward those two groups. The final inertia measures the remaining within-cluster spread.

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning