MATHEMATICAL FOUNDATIONS

THE MATHEMATICS BENEATH

Every model in this atlas stands on a handful of mathematical ideas. Each concept below answers one question first: why does this mathematics appear in machine learning at all?

00 Two learning paradigms

Supervised learning

Definition. Learn a mapping from features \(x\) to a known target \(y\) from labelled examples \((x_i, y_i)\).

Intuition. Studying with an answer key: every mistake is measurable, so the loss is well-defined.

Used by: all regression and classification models.

Unsupervised learning

Definition. Find structure in \(x\) alone — the objective is defined by the structure itself (compactness, density, variance).

Intuition. Sorting a drawer of mixed buttons without a picture of the finished sets.

Used by: K-Means, Hierarchical Clustering, DBSCAN, PCA.

01 Linear algebra

Scalars, vectors, matrices

Definition. A scalar is one number; a vector \(x \in \mathbb{R}^n\) is an ordered list; a matrix \(X \in \mathbb{R}^{n \times p}\) has one row per example, one column per feature.

Intuition. A dataset is a matrix; a single house is a vector of measurements.

\[ \hat{y} = X\beta \]

ML connection. The design matrix \(X\) and target \(y\) are the objects every regression model consumes.

Used by: Multiple Linear Regression, PCA, SVM.

Dot product

Definition. \(x^{\top}y = \sum_i x_i y_i\) — multiply pairwise, then sum.

Intuition. It measures alignment: large when vectors agree, zero when perpendicular.

\[ x^{\top}y = \sum_{i=1}^{n} x_i y_i = \|x\|\,\|y\|\cos\theta \]

ML connection. Every linear model's score \(z = \beta_0 + \beta^{\top}x\) is a dot product; kernels compute dot products in expanded spaces.

Used by: SVM, Logistic Regression.

Matrix multiplication

Definition. \((AB)_{ij} = \sum_k A_{ik}B_{kj}\) — row of \(A\) dotted with column of \(B\).

Intuition. Apply one transformation after another, to every row at once.

\[ (X^{\top}X)_{jk} = \sum_{i=1}^{n} x_{ij}\,x_{ik} \]

ML connection. \(X\beta\) predicts every target in one multiplication; \(X^{\top}X\) aggregates every feature co-variation into the Gram matrix.

Used by: Multiple Linear Regression, PCA, Ridge.

Norms

Definition. A norm measures a vector's length: \(L_2\) is the straight line, \(L_1\) the city block.

Intuition. L2 punishes big components hard; L1 taxes every unit equally — that is why L1 makes exact zeros.

\[ \|x\|_2=\sqrt{\textstyle\sum_i x_i^2}, \qquad \|x\|_1=\sum_i |x_i| \]

ML connection. Ridge penalises \(\|\beta\|_2^2\), Lasso penalises \(\|\beta\|_1\); K-Means minimises \(\|x-\mu\|_2^2\).

Used by: Ridge, Lasso, SVM, K-Means, KNN.

Eigenvalues & eigenvectors

Definition. \(v\) is an eigenvector of \(A\) with eigenvalue \(\lambda\) if \(Av = \lambda v\) — the matrix only stretches \(v\), never rotates it.

Intuition. The directions a transformation treats specially: the data's natural axes.

\[ \Sigma v_j = \lambda_j v_j \]

ML connection. PCA eigendecomposes the covariance matrix: eigenvectors are the principal directions, eigenvalues the variance along them.

Used by: PCA, Ridge (eigenvalue shift by \(\lambda\)).

02 Calculus

Derivatives

Definition. \(f'(x)\) is the instantaneous rate of change of \(f\) at \(x\).

Intuition. Stand on the loss surface: the derivative is the slope under your feet — it tells you which way is downhill.

\[ f'(x) = \lim_{h \to 0} \frac{f(x+h)-f(x)}{h} \]

ML connection. Setting \(\frac{dJ}{d\beta}=0\) solves ordinary least squares exactly; the sign of the slope dictates each gradient-descent step.

Used by: Linear Regression, Logistic Regression, Polynomial Regression.

Partial derivatives

Definition. \(\frac{\partial J}{\partial\beta_j}\) measures how \(J\) changes when only \(\beta_j\) moves.

Intuition. With many dials, sweep one dial at a time and watch the loss tick.

\[ \frac{\partial J}{\partial\beta_j} = \frac{1}{n}\sum_i -2(y_i - \beta^\top x_i)\,x_{ij} \]

ML connection. The gradient \(\nabla J\) is the vector of all partials — the compass gradient descent follows: \(\beta \leftarrow \beta - \eta\nabla J\).

Used by: Linear Regression, Logistic Regression, SVM, Ridge.

Gradient descent & the chain rule

Definition. Iteratively step opposite the gradient: \(\beta^{(t+1)} = \beta^{(t)} - \eta\nabla J(\beta^{(t)})\). The chain rule computes gradients through composed functions.

Intuition. Descend a foggy mountain by feeling only the slope underfoot; \(\eta\) is your stride — too small crawls, too large leaps across the valley.

\[ \beta \leftarrow \beta - \eta\,\nabla J(\beta) \]

ML connection. Gradient descent is the learning loop for almost every model that has no closed form — the chain rule distributes it through layered structures.

Used by: Logistic Regression, SVM, neural networks.

03 Probability

Probability & random variables

Definition. A probability \(P(A) \in [0,1]\) measures the chance of an outcome; a random variable assigns numbers to outcomes and comes with a distribution \(\mathrm{Pr}(X=x)\).

Intuition. The world is noisy — every label carries uncertainty. Probability is the calculus of that uncertainty.

\[ \sum_{x}\mathrm{Pr}(X=x) = 1 \]

ML connection. Classification models output probability estimates; uncertainty feeds risk decisions.

Used by: Naive Bayes, Logistic Regression, Decision Trees.

Conditional probability

Definition. \(P(A \mid B) = \frac{P(A \cap B)}{P(B)}\) is the chance of \(A\) given \(B\) is known.

Intuition. New evidence rescales what you believe: knowing it's raining changes the probability you'll need your umbrella.

\[ P(A \mid B) = \frac{P(A \cap B)}{P(B)} \]

ML connection. \(P(\text{class} \mid \text{features})\) is precisely the conditional probability every classifier estimates.

Used by: Naive Bayes, Logistic Regression.

Bayes' theorem

Definition. Flips a conditional probability: \(P(C\mid x) = \frac{P(x\mid C)\,P(C)}{P(x)}\).

Intuition. Start with a prior belief \(P(C)\), multiply by the likelihood of the evidence \(P(x\mid C)\), renormalise.

\[ P(C\mid x) = \frac{P(x\mid C)\,P(C)}{P(x)} \]

ML connection. Naive Bayes applies it with independence \(P(x\mid C)=\prod_j P(x_j \mid C)\) so the product stays computable.

Used by: Naive Bayes.

Distributions

Definition. A distribution describes which values a random variable takes and how often — the normal, Bernoulli and multinomial are the workhorses.

Intuition. The Gaussian models additive noise; the Bernoulli models a coin flip with probability \(p\).

\[ \mathcal{N}(x;\mu,\sigma^2)=\frac{1}{\sigma\sqrt{2\pi}}\exp\left(-\frac{(x-\mu)^2}{2\sigma^2}\right) \]

ML connection. Regression's noise is the Gaussian; classification's label is Bernoulli; text counts are multinomial.

Used by: Linear Regression, Naive Bayes, Logistic Regression.

04 Statistics

Mean, variance & standard deviation

Definition. \(\mu = \frac{1}{n}\sum_i x_i\) is the centre; \(\sigma^2 = \frac{1}{n}\sum_i (x_i-\mu)^2\) is the spread.

Intuition. Where is the cloud, and how wide is it?

\[ \mu=\frac{1}{n}\sum_i x_i,\qquad \sigma^2=\frac{1}{n}\sum_{i}(x_i-\mu)^2 \]

ML connection. Standardisation \((x-\mu)/\sigma\) makes features comparable; variance is what regression and PCA are built around.

Used by: Linear Regression, PCA, K-Means, KNN.

Covariance

Definition. How two features vary together: positive when they rise together, negative when they move oppositely.

Intuition. Height and weight covary positive; training time and error covary negative.

\[ \mathrm{Cov}(x,y)=\frac{1}{n}\sum_i (x_i-\bar{x})(y_i-\bar{y}) \]

ML connection. The covariance matrix \(\Sigma\) collects all pairwise covariances — the object PCA eigendecomposes and where multicollinearity lives.

Used by: PCA, Multiple Regression, Ridge.

Correlation

Definition. Covariance rescaled to \([-1,1]\): \(\rho_{xy}=\frac{\mathrm{Cov}(x,y)}{\sigma_x\sigma_y}\).

Intuition. Removes units, so correlation is comparable across very different features.

\[ \rho_{xy}=\frac{\mathrm{Cov}(x,y)}{\sigma_x\,\sigma_y} \in [-1,1] \]

ML connection. Strongly correlated features threaten a stable regression fit (multicollinearity) and are exactly what PCA compresses away.

Used by: Linear Regression, Ridge, PCA.

05 Optimization

Objective, loss & cost functions

Definition. A loss \(L(y,\hat{y})\) scores one prediction; the cost \(J=\frac{1}{n}\sum_i L(y_i,\hat{y}_i)\) averages the loss over the dataset. Training is minimising \(J\).

Intuition. The cost is a trainer: the model keeps the parameters that make it as small as possible.

\[ J(\theta)=\frac{1}{n}\sum_{i=1}^{n} L\big(y_i,\, f_\theta(x_i)\big) \]

ML connection. Squared error drives regression; log-loss drives classifiers; inertia drives K-Means.

Used by: all models.

Local vs global minimum

Definition. A global minimum is the lowest point of the cost surface; a local minimum is a valley that is not the lowest overall.

Intuition. In rolling terrain, gradient descent can settle in the first valley it finds even if a deeper one exists elsewhere.

\[ \theta^* = \arg\min_{\theta} J(\theta) \]

ML connection. Linear and logistic regression have convex bowls — one global minimum; K-Means and deep nets pay a local-minima tax.

Used by: Linear Regression, Logistic Regression, K-Means.

Regularization

Definition. An extra penalty that prefers simpler parameters: \(J_\lambda(\theta)=J(\theta)+\lambda \|\theta\|\).

Intuition. Complexity costs extra — so a coefficient is only kept when the data genuinely earn it.

\[ J_{\lambda}(\theta)=J(\theta)+\lambda \|\theta\| \]

ML connection. L1 (Lasso) drives coefficients to zero for feature selection; L2 (Ridge) shrinks them to tame collinearity. Both curb overfitting.

Used by: Ridge, Lasso, SVM.