Random Forests
Leo Breiman · Machine Learning
doi:10.1023/A:1010933404324
In short
Breiman grows many decision trees, each on a random resample of the data and choosing among a random subset of features at every split, then lets them vote. The randomness keeps the trees from making the same mistakes, so the forest is far more accurate and stable than any single tree.
Why it matters
It is still the first thing to try on tabular data, and the clearest demonstration that averaging diverse weak models beats one strong one.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 9 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Supervised Learning ✓ understood
Learning from examples paired with the correct answer, so a model can predict answers for new inputs it hasn't seen.
- Classification ✓ understood
A supervised learning task where the model assigns each input to one of a fixed set of categories, such as spam or not spam.
- Regression ✓ understood
A supervised learning task where the model predicts continuous numerical values rather than discrete categories.
- Decision Tree · read first ✓ understood
A tree-structured model that makes decisions by splitting data based on feature values, interpretable but prone to overfitting.
- Ensemble Learning · read first ✓ understood
Combining multiple models to produce better predictions than any individual model (bagging, boosting, stacking).
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training Data ✓ understood
The examples a model learns its weights from, kept separate from the validation and test data used to check how well it generalizes.
- Overfitting · read first ✓ understood
When a model fits its training data too closely, noise included, so it scores well on examples it has seen and poorly on new ones.