MACHINE LEARNING / ENSEMBLE

RANDOM FOREST

Many diverse decision trees vote together to reduce variance and improve robustness.

SupervisedEnsembleTree-based
Saved only in this browser

01 Overview

02 The Problem

A single decision tree can overfit one sample of the data. Random Forest reduces that instability by training many varied trees and combining their predictions.

03 Why It Matters

Bootstrap sampling and random feature selection decorrelate the trees, so averaging their errors usually produces a more stable model than one deep tree.

04 Intuition

Ask several imperfect trees the same question. Their different mistakes tend to cancel, while the shared signal survives the vote.

05 Mathematical Foundation

Each tree sees a bootstrap sample and a random subset of features at each split. Classification averages votes; regression averages numeric outputs.

06 The Equation

\[\hat{f}(x)=\frac{1}{B}\sum_{b=1}^{B}T_b(x)\]
  • \(B\) number of trees
  • \(T_b\) prediction from tree b
  • \(\hat f(x)\) ensemble prediction

07 How It Learns

  1. Draw a bootstrap sample.
  2. Grow a tree using random feature subsets.
  3. Repeat for many trees.
  4. Aggregate their predictions.

08 Algorithm

Bootstrap rows
↓
Grow varied trees
↓
Vote or average

09 Visual Explanation

A quick visual summary of how this model sees data and makes its prediction.

10 Worked Example

If four of five trees vote approve, the forest returns approve. A new bootstrap draw can change individual trees while leaving the majority stable.

11 Data & Features

12 Evaluation

13 Strengths

14 Limitations

15 When to Use

16 When Not to Use

17 Real-World Applications

19 60-Second Recap

20 Continue Learning