Start with data
Past house sales pair size, location and price. The target is a number, not a category.
Finding the straight-line relationship hidden inside data — the line that best explains a continuous target.
Linear Regression finds the line that best connects an input to a continuous outcome. It turns a cloud of observations into a relationship we can explain, evaluate and use.
Past house sales pair size, location and price. The target is a number, not a category.
Points scatter, but they can still tilt. A slope describes the average change in the target.
Each coefficient has a readable meaning: how much the prediction changes per unit of input.
Move the parameters and watch the line, residuals and error respond immediately.
Observe paired values.
Choose β₀ and β₁.
Compare y to ŷ.
Minimize total error.
Linear Regression searches for the line that minimizes prediction error across the observed data.
One line. Four ideas. Select a symbol to focus its meaning.
A residual is \(e_i = y_i - \hat{y}_i\). Least squares minimizes \(J = \frac{1}{n}\sum_i (y_i-\hat{y}_i)^2\), producing a smooth convex loss surface. For a full matrix, the normal equation is \(\beta = (X^T X)^{-1}X^T y\). Gradient descent follows \(\beta \leftarrow \beta - \eta\nabla J\) until the loss stops improving.
Centre x and y, compute \(s_{xx}=\sum(x_i-\bar{x})^2\) and \(s_{xy}=\sum(x_i-\bar{x})(y_i-\bar{y})\). Then \(\beta_1=s_{xy}/s_{xx}\) and \(\beta_0=\bar{y}-\beta_1\bar{x}\). Squaring the residuals makes the optimum unique for a non-degenerate dataset.
Start with β₀ = β₁ = 0, compute the gradient of the mean squared error, step downhill with learning rate η, and repeat until the loss stabilizes. The normal equation reaches the same optimum directly when the matrix is well-conditioned.
| x | y |
|---|---|
| 1 | 1 |
| 2 | 3 |
| 3 | 2 |
| 4 | 5 |
| 5 | 5 |
✓ The relationship is roughly linear
✓ A continuous prediction is needed
✓ Interpretability matters
✓ A fast baseline is useful
× The relationship is highly nonlinear
× Extreme outliers dominate
× Complex interactions matter
× Assumptions are seriously violated
Remember: Linear Regression learns the straight-line function that best represents the relationship between input and output variables.