THE MATHEMATICS BENEATH
Every model in this atlas stands on a handful of mathematical ideas. Each concept below answers one question first: why does this mathematics appear in machine learning at all?
00 Two learning paradigms
Supervised learning
Definition. Learn a mapping from features \(x\) to a known target \(y\) from labelled examples \((x_i, y_i)\).
Intuition. Studying with an answer key: every mistake is measurable, so the loss is well-defined.
Used by: all regression and classification models.
Unsupervised learning
Definition. Find structure in \(x\) alone — the objective is defined by the structure itself (compactness, density, variance).
Intuition. Sorting a drawer of mixed buttons without a picture of the finished sets.
Used by: K-Means, Hierarchical Clustering, DBSCAN, PCA.
01 Linear algebra
Scalars, vectors, matrices
Definition. A scalar is one number; a vector \(x \in \mathbb{R}^n\) is an ordered list; a matrix \(X \in \mathbb{R}^{n \times p}\) has one row per example, one column per feature.
Intuition. A dataset is a matrix; a single house is a vector of measurements.
ML connection. The design matrix \(X\) and target \(y\) are the objects every regression model consumes.
Used by: Multiple Linear Regression, PCA, SVM.
Dot product
Definition. \(x^{\top}y = \sum_i x_i y_i\) — multiply pairwise, then sum.
Intuition. It measures alignment: large when vectors agree, zero when perpendicular.
ML connection. Every linear model's score \(z = \beta_0 + \beta^{\top}x\) is a dot product; kernels compute dot products in expanded spaces.
Used by: SVM, Logistic Regression.
Matrix multiplication
Definition. \((AB)_{ij} = \sum_k A_{ik}B_{kj}\) — row of \(A\) dotted with column of \(B\).
Intuition. Apply one transformation after another, to every row at once.
ML connection. \(X\beta\) predicts every target in one multiplication; \(X^{\top}X\) aggregates every feature co-variation into the Gram matrix.
Used by: Multiple Linear Regression, PCA, Ridge.
Norms
Definition. A norm measures a vector's length: \(L_2\) is the straight line, \(L_1\) the city block.
Intuition. L2 punishes big components hard; L1 taxes every unit equally — that is why L1 makes exact zeros.
ML connection. Ridge penalises \(\|\beta\|_2^2\), Lasso penalises \(\|\beta\|_1\); K-Means minimises \(\|x-\mu\|_2^2\).
Eigenvalues & eigenvectors
Definition. \(v\) is an eigenvector of \(A\) with eigenvalue \(\lambda\) if \(Av = \lambda v\) — the matrix only stretches \(v\), never rotates it.
Intuition. The directions a transformation treats specially: the data's natural axes.
ML connection. PCA eigendecomposes the covariance matrix: eigenvectors are the principal directions, eigenvalues the variance along them.
02 Calculus
Derivatives
Definition. \(f'(x)\) is the instantaneous rate of change of \(f\) at \(x\).
Intuition. Stand on the loss surface: the derivative is the slope under your feet — it tells you which way is downhill.
ML connection. Setting \(\frac{dJ}{d\beta}=0\) solves ordinary least squares exactly; the sign of the slope dictates each gradient-descent step.
Used by: Linear Regression, Logistic Regression, Polynomial Regression.
Partial derivatives
Definition. \(\frac{\partial J}{\partial\beta_j}\) measures how \(J\) changes when only \(\beta_j\) moves.
Intuition. With many dials, sweep one dial at a time and watch the loss tick.
ML connection. The gradient \(\nabla J\) is the vector of all partials — the compass gradient descent follows: \(\beta \leftarrow \beta - \eta\nabla J\).
Used by: Linear Regression, Logistic Regression, SVM, Ridge.
Gradient descent & the chain rule
Definition. Iteratively step opposite the gradient: \(\beta^{(t+1)} = \beta^{(t)} - \eta\nabla J(\beta^{(t)})\). The chain rule computes gradients through composed functions.
Intuition. Descend a foggy mountain by feeling only the slope underfoot; \(\eta\) is your stride — too small crawls, too large leaps across the valley.
ML connection. Gradient descent is the learning loop for almost every model that has no closed form — the chain rule distributes it through layered structures.
Used by: Logistic Regression, SVM, neural networks.
03 Probability
Probability & random variables
Definition. A probability \(P(A) \in [0,1]\) measures the chance of an outcome; a random variable assigns numbers to outcomes and comes with a distribution \(\mathrm{Pr}(X=x)\).
Intuition. The world is noisy — every label carries uncertainty. Probability is the calculus of that uncertainty.
ML connection. Classification models output probability estimates; uncertainty feeds risk decisions.
Used by: Naive Bayes, Logistic Regression, Decision Trees.
Conditional probability
Definition. \(P(A \mid B) = \frac{P(A \cap B)}{P(B)}\) is the chance of \(A\) given \(B\) is known.
Intuition. New evidence rescales what you believe: knowing it's raining changes the probability you'll need your umbrella.
ML connection. \(P(\text{class} \mid \text{features})\) is precisely the conditional probability every classifier estimates.
Used by: Naive Bayes, Logistic Regression.
Bayes' theorem
Definition. Flips a conditional probability: \(P(C\mid x) = \frac{P(x\mid C)\,P(C)}{P(x)}\).
Intuition. Start with a prior belief \(P(C)\), multiply by the likelihood of the evidence \(P(x\mid C)\), renormalise.
ML connection. Naive Bayes applies it with independence \(P(x\mid C)=\prod_j P(x_j \mid C)\) so the product stays computable.
Used by: Naive Bayes.
Distributions
Definition. A distribution describes which values a random variable takes and how often — the normal, Bernoulli and multinomial are the workhorses.
Intuition. The Gaussian models additive noise; the Bernoulli models a coin flip with probability \(p\).
ML connection. Regression's noise is the Gaussian; classification's label is Bernoulli; text counts are multinomial.
Used by: Linear Regression, Naive Bayes, Logistic Regression.
04 Statistics
Mean, variance & standard deviation
Definition. \(\mu = \frac{1}{n}\sum_i x_i\) is the centre; \(\sigma^2 = \frac{1}{n}\sum_i (x_i-\mu)^2\) is the spread.
Intuition. Where is the cloud, and how wide is it?
ML connection. Standardisation \((x-\mu)/\sigma\) makes features comparable; variance is what regression and PCA are built around.
Used by: Linear Regression, PCA, K-Means, KNN.
Covariance
Definition. How two features vary together: positive when they rise together, negative when they move oppositely.
Intuition. Height and weight covary positive; training time and error covary negative.
ML connection. The covariance matrix \(\Sigma\) collects all pairwise covariances — the object PCA eigendecomposes and where multicollinearity lives.
Used by: PCA, Multiple Regression, Ridge.
Correlation
Definition. Covariance rescaled to \([-1,1]\): \(\rho_{xy}=\frac{\mathrm{Cov}(x,y)}{\sigma_x\sigma_y}\).
Intuition. Removes units, so correlation is comparable across very different features.
ML connection. Strongly correlated features threaten a stable regression fit (multicollinearity) and are exactly what PCA compresses away.
Used by: Linear Regression, Ridge, PCA.
05 Optimization
Objective, loss & cost functions
Definition. A loss \(L(y,\hat{y})\) scores one prediction; the cost \(J=\frac{1}{n}\sum_i L(y_i,\hat{y}_i)\) averages the loss over the dataset. Training is minimising \(J\).
Intuition. The cost is a trainer: the model keeps the parameters that make it as small as possible.
ML connection. Squared error drives regression; log-loss drives classifiers; inertia drives K-Means.
Used by: all models.
Local vs global minimum
Definition. A global minimum is the lowest point of the cost surface; a local minimum is a valley that is not the lowest overall.
Intuition. In rolling terrain, gradient descent can settle in the first valley it finds even if a deeper one exists elsewhere.
ML connection. Linear and logistic regression have convex bowls — one global minimum; K-Means and deep nets pay a local-minima tax.
Used by: Linear Regression, Logistic Regression, K-Means.
Regularization
Definition. An extra penalty that prefers simpler parameters: \(J_\lambda(\theta)=J(\theta)+\lambda \|\theta\|\).
Intuition. Complexity costs extra — so a coefficient is only kept when the data genuinely earn it.
ML connection. L1 (Lasso) drives coefficients to zero for feature selection; L2 (Ridge) shrinks them to tame collinearity. Both curb overfitting.