Learning Representations by Back-Propagating Errors
David E. Rumelhart et al. · Nature
doi:10.1038/323533a0
In short
The paper shows how to compute, layer by layer from the output backwards, how much each weight in a multi-layer network contributed to the error, and nudge it accordingly. Trained this way, hidden units learn useful internal features on their own.
Why it matters
Backpropagation is how essentially every neural network, from LeNet to today’s LLMs, is trained.
Read first
The 4 Field Guide ideas this paper leans on.
Starting from scratch? The full route 7 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network · read first ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Loss Function · read first ✓ understood
A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.
- Gradient Descent · read first ✓ understood
An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.
- Backpropagation · read first ✓ understood
The algorithm for computing gradients of the loss with respect to network weights, enabling training through gradient descent.