K-MEANS CLUSTERING
Group unlabeled points around K moving centroids until assignments stop changing.
01 Overview
02 The Problem
Suppose customer records have no labels. K-Means groups nearby records into a chosen number of segments by repeatedly assigning points to their nearest centre and moving each centre to its group mean.
03 Why It Matters
Clustering turns an unlabelled dataset into a first map of its structure. It is useful for segmentation, compression, anomaly review, and exploratory analysis.
04 Intuition
Place k movable centres on the map. Every point joins its closest centre, then every centre moves to the average of its members. Alternating those two steps settles into a locally good partition.
05 Mathematical Foundation
The objective is within-cluster squared distance. Euclidean distance makes the mean the centre that minimises the total squared error for its assigned points.
06 The Equation
- \(C_k\) points assigned to cluster k
- \(\mu_k\) mean of cluster k
- \(J\) total within-cluster inertia
07 How It Learns
- Choose k initial centres.
- Assign each point to its nearest centre.
- Replace each centre with its assigned mean.
- Repeat until assignments stop changing.
08 Algorithm
09 Visual Explanation
A quick visual summary of how this model sees data and makes its prediction.
10 Worked Example
With k = 2, points near (1, 1) and (8, 8) quickly pull their centres toward those two groups. The final inertia measures the remaining within-cluster spread.