MACHINE LEARNING / CLASSIFICATION

NAIVE BAYES

Bayes theorem at full speed — powered by a deliberately strong conditional-independence assumption.

SupervisedClassificationProbabilistic
Saved only in this browser

01 Overview

02 The Problem

A mail service reads one email and must decide: spam or not? It cannot read prose, but it can see which words fired. Alone each word is fragile — “free” screams spam, “meeting” screams ham — yet combined they should be decisive. The problem is fusing many weak, independent clues into one probability.

03 Why It Matters

Bayes' theorem reverses cause and effect: from \(P(\text{word}\mid\text{spam})\) to \(P(\text{spam}\mid\text{words})\). It is the honest way to start from a prior and update with evidence; speed and accuracy on word counts made it the text-champion for decades.

04 Intuition

Every word votes, weighted by how diagnostic it is. “Free” appears in spam 55% of the time but ham only 4%, so seeing it pushes strongly toward spam. The naive assumption treats words as conditionally independent given the class — a fiction, but powerful when no single word is perfect.

05 Mathematical Foundation

\(P(C_k\mid x)\propto P(C_k)\prod_j P(x_j\mid C_k)\). Products underflow, so we sum in log space: \(\log P=\log P(C_k)+\sum_j \log P(x_j\mid C_k)\). Laplace smoothing \((+\alpha)\) dodges zero probabilities when a word is absent in training.

06 The Equation

\[ P(\text{spam}\mid \mathbf{w}) = \frac{P(\text{spam})\prod_{j} P(w_j\mid\text{spam})}{P(\text{spam})\prod_{j}P(w_j\mid\text{spam})+P(\text{ham})\prod_{j}P(w_j\mid\text{ham})} \]
  • \(P(\text{spam})\) prior probability of spam in the population
  • \(P(w_j\mid\text{spam})\) likelihood of word \(w_j\) in a spam email
  • \(\prod_j\) the naive conditional-independence product over words
  • 07 How It Learns

    1. Count how often each word appears in spam vs ham emails.
    2. Estimate \(P(w_j\mid C_k)\) with Laplace smoothing to dodge zeros.
    3. Fix a prior \(P(\text{spam})\) — adjust the slider for your inbox.
    4. At predict: apply Bayes in log space and threshold.

    08 Algorithm

    Training counts
    ↓
    Smoothed P(word|class)
    ↓
    Bayes rule in log space
    ↓
    Sum evidence → posterior
    ↓
    Threshold → label

    09 Visual Explanation

    A quick visual summary of how this model sees data and makes its prediction.

    10 Worked Example

    Email contains “free” (P=0.55) and “winner” (0.40); absent “meeting” (so factor 0.70). Prior P(spam)=0.4. Log-odds = ln(0.4/0.6)+ln(0.55/0.04)+ln(0.40/0.02)+ln(0.70/0.30) ≈ 0.56 → P(spam) ≈ 0.64. Positive words swung it from 0.40 to 0.64 — Bayes in one paragraph.

    11 Data & Features

    12 Evaluation

    13 Strengths

    14 Limitations

    15 When to Use

    16 When Not to Use

    17 Real-World Applications

    19 60-Second Recap

    20 Continue Learning