MACHINE LEARNING / REGRESSION

LINEAR REGRESSION

Finding the straight-line relationship hidden inside data — the line that best explains a continuous target.

SupervisedRegressionLinear
Saved only in this browser

01 Overview

01 / UNDERSTAND

What problem does a straight line solve?

Linear Regression finds the line that best connects an input to a continuous outcome. It turns a cloud of observations into a relationship we can explain, evaluate and use.

01

Start with data

Past house sales pair size, location and price. The target is a number, not a category.

02

Look for a trend

Points scatter, but they can still tilt. A slope describes the average change in the target.

03

Make it explainable

Each coefficient has a readable meaning: how much the prediction changes per unit of input.

REAL-WORLD QUESTIONCan previous house sales teach us the mathematical relationship between size and price?
02 / VISUALIZE

See Linear Regression in Action

Move the parameters and watch the line, residuals and error respond immediately.

LIVE MODEL
Drag a point to change the data
03 / INTUITION

From scattered data to a useful line

01Data

Observe paired values.

→
02Guess a line

Choose β₀ and β₁.

→
03Measure error

Compare y to ŷ.

→
04Find best fit

Minimize total error.

Linear Regression searches for the line that minimizes prediction error across the observed data.

04 / MATHEMATICS

The Mathematics Behind Linear Regression

ŷ = β₀ + β₁x

One line. Four ideas. Select a symbol to focus its meaning.

Explore the Mathematics +

A residual is \(e_i = y_i - \hat{y}_i\). Least squares minimizes \(J = \frac{1}{n}\sum_i (y_i-\hat{y}_i)^2\), producing a smooth convex loss surface. For a full matrix, the normal equation is \(\beta = (X^T X)^{-1}X^T y\). Gradient descent follows \(\beta \leftarrow \beta - \eta\nabla J\) until the loss stops improving.

Show Full Derivation +

Centre x and y, compute \(s_{xx}=\sum(x_i-\bar{x})^2\) and \(s_{xy}=\sum(x_i-\bar{x})(y_i-\bar{y})\). Then \(\beta_1=s_{xy}/s_{xx}\) and \(\beta_0=\bar{y}-\beta_1\bar{x}\). Squaring the residuals makes the optimum unique for a non-degenerate dataset.

05 / TRAINING

How the model learns

01Guess parametersβ₀ = 0 · β₁ = 0
→
02Predictŷ = β₀ + β₁x
→
03Calculate errorMSE = 24.7
→
04UpdateStep downhill
→
05Best fitMSE ↓ 0.42
LOSS CURVEError falls as the line learns
See the training algorithm +

Start with β₀ = β₁ = 0, compute the gradient of the mean squared error, step downhill with learning rate η, and repeat until the loss stabilizes. The normal equation reaches the same optimum directly when the matrix is well-conditioned.

06 / PRACTICE

Work it through

WORKED EXAMPLE

Five observations, one equation

xy
11
23
32
45
55
  1. Average: x̄ = 3, ȳ = 3
  2. Slope: β₁ = sxy / sxx = 12 / 10 = 1.2
  3. Intercept: β₀ = ȳ − β₁x̄ = −0.6
ŷ = −0.6 + 1.2x For x = 6, predict ŷ = 6.6
07 / APPLY

Know when it earns its place

Great choice when

✓ The relationship is roughly linear

✓ A continuous prediction is needed

✓ Interpretability matters

✓ A fast baseline is useful

Avoid or be careful when

× The relationship is highly nonlinear

× Extreme outliers dominate

× Complex interactions matter

× Assumptions are seriously violated

⌂HousingSize → Price
↗MarketingSpend → Sales
◒EnergyTemperature → Consumption
✦EducationStudy hours → Score
60-SECOND RECAP
Data↓Assume relationship↓Fit line↓Minimize error↓Evaluate↓Predict

Remember: Linear Regression learns the straight-line function that best represents the relationship between input and output variables.

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning