MACHINE LEARNING / REGRESSION

MULTIPLE LINEAR REGRESSION

One target, many drivers — a hyperplane that weighs every feature to predict a scalar.

SupervisedRegressionLinear
Saved only in this browser

01 Overview

02 The Problem

A medical researcher predicts a patient's blood-pressure change from several measurements at once: age, weight, baseline BP, exercise minutes. No single number tells the story. Real targets rarely have one cause, so the question is “how much, given several inputs?” — and the answer must still say what each input is worth.

03 Why It Matters

Multiple regression is applied statistics' workhorse. Its power is ceteris paribus reading: a coefficient is a feature’s effect holding the others fixed — exactly how a doctor reasons (“age raises risk by X per year, independent of weight”).

04 Intuition

One feature tilts a line. Two features tilt a plane; \(p\) features define a hyperplane in \(p{+}1\) dimensions. Fitting rotates that sheet until the total squared vertical distance to every point is minimal — the ruler game generalised.

05 Mathematical Foundation

Stack examples in the design matrix \(X\) (one column of 1s for the intercept) and targets in \(y\). Predictions are \(\hat y = X\beta\) and the cost \(\|y-X\beta\|^2\) is convex and quadratic in \(\beta\). Setting the gradient \(-2X^{\top}(y-X\beta)\) to zero gives the normal equations \(X^{\top}X\,\beta = X^{\top}y\).

06 The Equation

\[ \hat{\beta} = (X^{\top}X)^{-1}X^{\top}y \]
  • \(X\) \(n \times (p{+}1)\) design matrix; first column all 1s
  • \(y\) vector of \(n\) observed targets
  • \(\beta\) coefficients \((\beta_0,\dots,\beta_p)\) being learned
  • \(X^{\top}X\) Gram matrix of feature co-variation
  • \(X^{\top}y\) cross-covariance of features with the target

07 How It Learns

  1. Assemble \(X\) and \(y\) — one row per example, intercept column first.
  2. Form \(X^{\top}X\) and \(X^{\top}y\) — sums of squares and cross-products.
  3. Gaussian elimination on \(X^{\top}X\,\beta = X^{\top}y\).
  4. Read coefficients. Each \(\beta_j\) is the marginal effect of feature \(j\).

08 Algorithm

Data table (n × p + target)
↓
Build \(X\) and \(y\)
↓
Compute \(X^{\top}X,\;X^{\top}y\)
↓
Solve → \(\hat\beta\)
↓
\(\hat{y}=x^{\top}\hat\beta\)

09 Visual Explanation

A quick visual summary of how this model sees data and makes its prediction.

10 Worked Example

Fitted \(\hat{y}=2+1.2\,x_1-0.8\,x_2\). Client \(x_1=40,x_2=5\): \(\hat{y}=2+48-4=46\). Age adds 1.2 per year; exercise removes 0.8 — and clients differing only in exercise by 2 score \(2\times0.8=1.6\), ceteris paribus in action.

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning