MACHINE LEARNING / REGRESSION

LASSO REGRESSION

The feature selector — an L1 penalty that can drive coefficients exactly to zero.

SupervisedRegressionRegularized
Saved only in this browser

01 Overview

02 The Problem

A bioinformatician has 20,000 gene expressions but only 100 patients. Most genes are noise; a handful drive the biomarker. Ordinary regression spreads credit thin across every gene. The problem is “which inputs matter?” — not just predicting, but picking the real drivers.

03 Why It Matters

Lasso adds an L1 penalty \(\lambda\|\!\|\beta\|\!\|\) which can drive coefficients exactly to zero. Where Ridge shrinks, Lasso selects — it returns a sparse model containing only the features it kept. That is interpretability with a built-in verdict.

04 Intuition

The L1 ball has sharp corners. As the penalty grows, the error contour first kisses a corner of that ball — and at a corner, some coordinates are exactly zero. So Lasso keeps a few strong features and discards the rest, automatically. It is a shrink-and-select operator.

05 Mathematical Foundation

Objective \(J_{Lasso}=\|y-X\beta\|^2+\lambda\|\beta\|_1\). The L1 term is not differentiable at zero, so there is no closed form — coordinate descent sweeps each \(\beta_j\) to its soft-thresholded optimum while holding the rest fixed, repeating until convergence.

06 The Equation

\[ \hat{\beta}_j \leftarrow S_{\lambda/2}\!\left(\beta_j + \frac{1}{n}\sum_{i} x_{ij}\,r_i\right), \quad S_{\lambda/2}(z)=\operatorname{sign}(z)\max(|z|-\lambda/2,\,0) \]
  • \(S_{\lambda/2}\) the soft-thresholding operator — shrinks toward 0, zeros the small ones
  • \(r_i\) partial residual for example \(i\)
  • \(\lambda\) strength of the L1 penalty
  • \(\|\beta\|_1\) sum of absolute coefficients — the sparsity-inducing cost

07 How It Learns

  1. Standardise all features.
  2. Cycle through one coefficient at a time.
  3. Soft-threshold each \(\beta_j\), driving small ones to 0.
  4. Repeat until coefficients settle; tune \(\lambda\) by CV.

08 Algorithm

Data (X, y)
↓
Standardise
↓
Coordinate descent loop
↓
Soft-threshold → β̂ (some exactly 0)
↓
ŷ = Xβ̂

09 Visual Explanation

A quick visual summary of how this model sees data and makes its prediction.

10 Worked Example

Predict house price from 6 features; only size, age, and location truly matter. At \(\lambda=0\) all six coefficients wiggle. At \(\lambda=1.4\) the tax-rate and school-rank coefficients hit zero — the model has chosen size, age, location and ignored the rest. Sparse, explainable, defensible.

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning