Long Short-Term Memory
Sepp Hochreiter et al. · Neural Computation
doi:10.1162/neco.1997.9.8.1735
In short
Recurrent networks forget quickly because error signals shrink or blow up as they flow back through time. The LSTM adds a memory cell with gates that decide what to write, keep and read, letting the gradient flow across hundreds of steps.
Why it matters
LSTMs powered speech recognition, translation and text generation for twenty years, until transformers took over.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 11 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Recurrent Neural Network · read first ✓ understood
A neural network architecture with loops that allow information to persist, designed for sequential data like text and time series.
- Activation Function ✓ understood
A non-linear function applied to neuron outputs that introduces non-linearity, enabling networks to learn complex patterns.
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Loss Function ✓ understood
A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.
- Gradient Descent ✓ understood
An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.
- Backpropagation ✓ understood
The algorithm for computing gradients of the loss with respect to network weights, enabling training through gradient descent.
- Vanishing Gradient · read first ✓ understood
A problem where gradients become extremely small during backpropagation, preventing deep networks from learning effectively.
- Long Short-Term Memory · read first ✓ understood
A type of RNN architecture with gates that can learn long-term dependencies, solving the vanishing gradient problem.