Landmark 3 stops to get here · leads to 11

Gradient Descent

An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.

Your route here

3 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Loss Function ✓ understood

    A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.

  4. Gradient Descent · you are here ✓ understood

Picture it

minimumstartEACH STEP MOVES DOWNHILL: θ ← θ − η · ∇L(θ)
Each step moves against the gradient, downhill across the contours, zig-zagging toward the loss minimum.

Gradient descent is how nearly every neural network finds good weights. Picture the loss as a landscape: each position is one setting of the model’s parameters, and height is how wrong the model is. You can’t see the whole landscape, but you can feel the slope underfoot. Gradient descent takes a step downhill, feels the slope again, and repeats.

The slope is the gradient: for each parameter, how much the loss would change if that parameter moved slightly. The update rule is one line. Subtract the gradient, scaled by a small number called the learning rate, from the current parameters. The method goes back to Cauchy in 1847. Today it runs on models with billions of parameters, with backpropagation computing the gradient.

The variants

The exact gradient requires running the entire dataset through the model for every single step. So in practice there are three options:

Smarter steps

Plain gradient descent struggles in long, narrow valleys, zig-zagging from wall to wall instead of moving along the floor. Momentum keeps a running average of past gradients, so steps build speed in directions that stay consistent and cancel out in ones that flip. Adaptive methods give each parameter its own step size. Adam, introduced in 2014, combines both ideas using running estimates of the gradient’s mean and variance, and it’s a common default for training deep networks.

The catch

The learning rate matters more than almost any other setting. Too large and the loss bounces around or blows up; too small and training crawls or stalls. Most large training runs change it over time, warming it up at the start and decaying it toward the end.

Gradient descent also only finds a low point near where it happens to be. For a convex problem, with a single bowl-shaped valley, that’s the best possible answer. Neural network loss landscapes aren’t convex, so there’s no such guarantee, although in practice the solutions it reaches in large networks usually work well.

Where it sits

Explore nearby

In the research

All papers →

3 papers that build on Gradient Descent .

Sources

  1. Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning . MIT Press, 2016; section 4.3, Gradient-Based Optimization
  2. Sebastian Ruder, "An overview of gradient descent optimization algorithms" . 2016
  3. Kingma and Ba, "Adam: A Method for Stochastic Optimization" . 2014; ICLR 2015