MACHINE LEARNING / REGRESSION

RIDGE REGRESSION

Least squares with an L2 safety net — shrink coefficients to gain stability under multicollinearity.

SupervisedRegressionRegularized
Saved only in this browser

01 Overview

02 The Problem

An econometrician forecasts next quarter’s growth from dozens of correlated indicators: GDP, inflation, unemployment, interest rates. Plain regression will fit, but the coefficients fly apart — a tiny data change swings them by 100%. The problem is multicollinearity: features that move together make the least-squares solution numerically unstable.

03 Why It Matters

Ridge regression trades a pinprick of bias for a huge drop in variance. When \(X^{\top}X\) is nearly singular, the ridge term \(\lambda I\) makes it invertible and the coefficients behave. It is the default fix whenever features outnumber samples or march in step.

04 Intuition

Ordinary least squares tries to explain every wiggle, even noise — so coefficients inflate to chase outliers. Ridge shrinks every coefficient toward zero (it never zeros them), damping the wobble. Raise \(\lambda\) and the fit grows calmer, at the price of a little bias. The bias–variance tradeoff made tangible.

05 Mathematical Foundation

The objective adds an L2 penalty: \(J_{Ridge}=\|y-X\beta\|^2+\lambda\|\beta\|^2\). Differentiating gives \((X^{\top}X+\lambda I)\beta=X^{\top}y\). The \(\lambda I\) diagonal stabilises the matrix: even if features are collinear, \(X^{\top}X+\lambda I\) stays invertible. Shrinkage is the price of stability.

06 The Equation

\[ \hat{\beta} = (X^{\top}X + \lambda I)^{-1}X^{\top}y \]
  • \(X\) \(n\times p\) design matrix of features
  • \(y\) vector of \(n\) observed targets
  • \(\lambda\) regularisation strength — 0 recovers OLS, ∞ shrinks everything to 0
  • \(I\) identity matrix; the L2 penalty added to the diagonal

07 How It Learns

  1. Standardise features so the penalty is fair to all.
  2. Add \(\lambda I\) to the Gram matrix \(X^{\top}X\).
  3. Solve \((X^{\top}X+\lambda I)\beta=X^{\top}y\).
  4. Tune \(\lambda\) by cross-validation — small is less biased, large is more stable.

08 Algorithm

Data (X, y)
↓
Scale features
↓
Add λI to XᵀX
↓
Solve → β̂
↓
ŷ = Xβ̂

09 Visual Explanation

A quick visual summary of how this model sees data and makes its prediction.

10 Worked Example

Forecast growth from 4 collinear indicators; OLS gives \(\beta=(12,-8,9,-11)\) — signs flip on reordering. Ridge with \(\lambda=4\) gives \(\beta=(1.9,0.6,-0.4,0.3)\): stable, sensible, every coefficient positive-ish. A little bias bought a lot of trustworthiness.

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning