NAIVE BAYES
Bayes theorem at full speed — powered by a deliberately strong conditional-independence assumption.
01 Overview
02 The Problem
A mail service reads one email and must decide: spam or not? It cannot read prose, but it can see which words fired. Alone each word is fragile — “free” screams spam, “meeting” screams ham — yet combined they should be decisive. The problem is fusing many weak, independent clues into one probability.
03 Why It Matters
Bayes' theorem reverses cause and effect: from \(P(\text{word}\mid\text{spam})\) to \(P(\text{spam}\mid\text{words})\). It is the honest way to start from a prior and update with evidence; speed and accuracy on word counts made it the text-champion for decades.
04 Intuition
Every word votes, weighted by how diagnostic it is. “Free” appears in spam 55% of the time but ham only 4%, so seeing it pushes strongly toward spam. The naive assumption treats words as conditionally independent given the class — a fiction, but powerful when no single word is perfect.
05 Mathematical Foundation
\(P(C_k\mid x)\propto P(C_k)\prod_j P(x_j\mid C_k)\). Products underflow, so we sum in log space: \(\log P=\log P(C_k)+\sum_j \log P(x_j\mid C_k)\). Laplace smoothing \((+\alpha)\) dodges zero probabilities when a word is absent in training.
06 The Equation
- \(P(\text{spam})\) prior probability of spam in the population
- \(P(w_j\mid\text{spam})\) likelihood of word \(w_j\) in a spam email
- \(\prod_j\) the naive conditional-independence product over words
- Count how often each word appears in spam vs ham emails.
- Estimate \(P(w_j\mid C_k)\) with Laplace smoothing to dodge zeros.
- Fix a prior \(P(\text{spam})\) — adjust the slider for your inbox.
- At predict: apply Bayes in log space and threshold.
07 How It Learns
08 Algorithm
09 Visual Explanation
A quick visual summary of how this model sees data and makes its prediction.
10 Worked Example
Email contains “free” (P=0.55) and “winner” (0.40); absent “meeting” (so factor 0.70). Prior P(spam)=0.4. Log-odds = ln(0.4/0.6)+ln(0.55/0.04)+ln(0.40/0.02)+ln(0.70/0.30) ≈ 0.56 → P(spam) ≈ 0.64. Positive words swung it from 0.40 to 0.64 — Bayes in one paragraph.