Algorithms · 2001
Random Forests
Grow hundreds of mediocre decision trees, let them vote, and the crowd beats any single expert.
In 2001, Berkeley statistician Leo Breiman published 'Random Forests', a method built on a counterintuitive bet: many weak, unstable models, combined, can be more accurate and more robust than one carefully tuned model.
A single decision tree is easy to read but notoriously twitchy; change a few training examples and it can grow completely differently. Breiman turned that instability into an asset by training hundreds of trees, each on a random bootstrap sample of the data and each allowed to consider only a random subset of features at every split.
The randomness ensures the trees make different mistakes. When they vote, or average, their errors partly cancel while their shared signal reinforces. The technique stood on his own earlier work on bagging and on Tin Kam Ho's random subspace method.
The payoff was practical. Random forests need little tuning, resist overfitting, handle thousands of features, and come with a free estimate of which variables matter most. For years they were the go-to method for tabular data in competitions and industry alike.
They remain a workhorse. Whenever a bank scores a loan or a hospital flags a risk from spreadsheet-style data, there is a good chance a forest, or its gradient-boosted cousin, is doing the work.
Breiman, who died in 2005, spent his late career arguing that predictive accuracy deserved as much respect as tidy statistical models. Random forests were his most persuasive argument.
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.