Gradient descent is how nearly every neural network finds good weights. Picture the loss as a landscape: each position is one setting of the model’s parameters, and height is how wrong the model is. You can’t see the whole landscape, but you can feel the slope underfoot. Gradient descent takes a step downhill, feels the slope again, and repeats.
The slope is the gradient: for each parameter, how much the loss would change if that parameter moved slightly. The update rule is one line. Subtract the gradient, scaled by a small number called the learning rate, from the current parameters. The method goes back to Cauchy in 1847. Today it runs on models with billions of parameters, with backpropagation computing the gradient.
The variants
The exact gradient requires running the entire dataset through the model for every single step. So in practice there are three options:
- Batch gradient descent: the whole dataset per step. Precise, but slow and memory-hungry on large data.
- Stochastic gradient descent: one example per step. Cheap, but each step is noisy.
- Mini-batch gradient descent: a small batch per step. This is the default, and it’s usually what people mean when they say SGD.
Smarter steps
Plain gradient descent struggles in long, narrow valleys, zig-zagging from wall to wall instead of moving along the floor. Momentum keeps a running average of past gradients, so steps build speed in directions that stay consistent and cancel out in ones that flip. Adaptive methods give each parameter its own step size. Adam, introduced in 2014, combines both ideas using running estimates of the gradient’s mean and variance, and it’s a common default for training deep networks.
The catch
The learning rate matters more than almost any other setting. Too large and the loss bounces around or blows up; too small and training crawls or stalls. Most large training runs change it over time, warming it up at the start and decaying it toward the end.
Gradient descent also only finds a low point near where it happens to be. For a convex problem, with a single bowl-shaped valley, that’s the best possible answer. Neural network loss landscapes aren’t convex, so there’s no such guarantee, although in practice the solutions it reaches in large networks usually work well.