MACHINE LEARNING / REGRESSION

POLYNOMIAL REGRESSION

When a straight line cannot bend enough — fit curves while keeping the least-squares optimality.

SupervisedRegressionNon-linear
Saved only in this browser

01 Overview

02 The Problem

DNA grows along a curve, not a line. A drug's dose-response bends; a child's height surges then stops. The straight line of Linear Regression cannot bend, so it underfits systematically — the residual still holds the shape. The problem is model non-linearity while keeping regression interpretable.

03 Why It Matters

Almost no natural relationship is globally linear. Polynomial regression lets a single feature bend the fit — a dose–response curve, a growth curve, a calibration curve — while reusing the fast, closed-form least-squares machinery.

04 Intuition

It is still a linear model — but the feature is expanded into powers. Treat \(x, x^2, x^3, \dots\) as new columns; the fit stays linear in the parameters, so the bowl is still smooth and has one bottom. The degree \(n\) is the flexibility dial: too low underfits, too high overfits and extrapolates wildly.

05 Mathematical Foundation

Design matrix \(X=[1, x, x^2, \dots, x^n]\). The cost \(\|y-X\beta\|^2\) is still quadratic in \(\beta\), so the normal equations \(\beta=(X^{\top}X)^{-1}X^{\top}y\) still apply. The basis functions \(\phi_j(x)=x^j\) turn a non-linear shape into a linear-algebra problem.

06 The Equation

\[ \hat{y} = \beta_0 + \beta_1 x + \beta_2 x^2 + \cdots + \beta_n x^n = \sum_{j=0}^{n}\beta_j\, x^{j} \]
  • \(n\) polynomial degree — the flexibility dial
  • \(\beta_j\) coefficient of the j-th power
  • \(x^j\) the j-th basis function of \(x\)
  • \(\hat{y}\) predicted continuous output

07 How It Learns

  1. Basis-expand each input into \([1, x, x^2, \dots, x^n]\).
  2. Solve normal equations for \(\beta\) (still one linear-algebra step).
  3. Diagnose degree \(n\) — bias/variance; pick via cross-validation.
  4. Predict \(\sum \beta_j x^{j}\) for new \(x\).

08 Algorithm

Data \((x_i,y_i)\)
↓
Expand to \(x^j\) basis
↓
OLS → \(\beta_j\)
↓
\(\hat{y}=\sum\beta_j x^j\)

09 Visual Explanation

A quick visual summary of how this model sees data and makes its prediction.

10 Worked Example

Dose \(x\) vs efficacy \(y\). Points suggest a plateau near \(x=8\). A degree-1 line is wrong; degree-5 fits but swings to \(\hat{y}=-2\) at \(x=10\) (impossible). Degree-2 \(\hat{y}=0.1+1.0x-0.08x^2\) plateaus — a parsimonious, sensible curve. This is the bias–variance dance in one number: \(n\).

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning