MACHINE LEARNING / CLASSIFICATION

LOGISTIC REGRESSION

A linear model that speaks in probabilities — the sigmoid turns scores into class predictions.

SupervisedClassificationLinear
Saved only in this browser

01 Overview

02 The Problem

A doctor must answer yes/no: does this tumour be malignant? The evidence is a number — a blood-marker concentration — not a category label she already has. The problem is turning a continuous measurement into a calibrated probability of the positive class, then crossing a threshold that balances missed cancers against false alarms.

03 Why It Matters

Malignancy is binary in outcome but probabilistic in evidence. Logistic regression gives a probability, not a hard label — so a clinician can pick the threshold that suits the cost of a false negative versus a false positive. That calibration is what makes it the default for ranking and for risk scoring.

04 Intuition

Linear regression on a 0/1 label can overshoot — predicting "1.3" or "-0.2", which is nonsense as a probability. The sigmoid squashes any real score into (0,1): small scores stay near 0, large ones saturate near 1, and the transition is smooth. A single decision threshold turns that probability into a label.

05 Mathematical Foundation

A linear score \(z=\beta_0+\beta^{\top}x\) is passed through the logistic function \(\sigma(z)=1/(1+e^{-z})\). We fit \(\beta\) by maximum likelihood — equivalently by minimising the log-loss (cross-entropy) \(-\sum[y_i\log p_i+(1-y_i)\log(1-p_i)]\), which is convex so the optimum is unique.

06 The Equation

\[ p = \sigma(z)=\frac{1}{1+e^{-z}}, \quad z=\beta_0+\beta^{\top}x, \quad \mathcal{L}=-\sum_{i}\bigl[y_i\log p_i+(1-y_i)\log(1-p_i)\bigr] \]
  • \(z\) the linear score from the features
  • \(\sigma\) the sigmoid — maps any real \(z\) into (0,1)
  • \(p\) predicted probability of the positive class
  • \(\mathcal{L}\) log-loss — the objective minimised to fit

07 How It Learns

  1. Start flat. \(\beta_0=\beta_1=0\) ⇒ probability 0.5 everywhere.
  2. Compute gradient of log-loss (it is nonzero until separation is reached).
  3. Step downhill \(\beta\leftarrow\beta-\eta\,\nabla\mathcal{L}\).
  4. Read off probabilities via the sigmoid and cross one threshold.

08 Algorithm

Labelled data (xᵢ, yᵢ∈{0,1})
↓
Score z=β₀+βᵀx
↓
σ(z) → p
↓
Gradient ↓ log-loss
↓
Threshold p → label

09 Visual Explanation

A quick visual summary of how this model sees data and makes its prediction.

10 Worked Example

Marker \(x\): negatives avg 2.0, positives avg 6.2. Fit \(\beta_0=-9.5,\beta_1=2.7\). For \(x=5\): \(z=-9.5+13.5=4\), \(p=1/(1+e^{-4})=0.98\). At threshold 0.5 → “malignant”. Lower threshold to 0.3 ⇒ catches earlier cases but raises false alarms — the threshold is the clinical policy.

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning