Training Dec 2014

Adam: A Method for Stochastic Optimization

Diederik P. Kingma et al. · ICLR 2015

arXiv:1412.6980

In short

Adam gives every parameter its own step size, using running averages of its recent gradients and of their squares. It is simple, cheap in memory, copes with noisy and sparse gradients, and usually works well with little tuning.

Why it matters

Adam (and its descendant AdamW) is the default optimizer for training almost every modern network, LLMs included.

Read first

The 4 Field Guide ideas this paper leans on.

Starting from scratch? The full route 7 ideas · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Loss Function ✓ understood

    A function that scores how wrong a model's prediction is as a single number, which training then works to make as small as possible.

  4. Gradient Descent · read first ✓ understood

    An optimization method that repeatedly moves a model's parameters a small step in the direction that most reduces the loss.

  5. Learning Rate · read first ✓ understood

    A hyperparameter controlling the step size in gradient descent - too high causes instability, too low slows convergence.

  6. Momentum · read first ✓ understood

    An optimization technique that accelerates gradient descent by accumulating past gradients, helping escape local minima.

  7. Adam Optimizer · read first ✓ understood

    An adaptive learning rate optimization algorithm combining momentum and RMSprop, widely used for training neural networks.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers