Landmark 1 stop to get here · leads to 15

Training

The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

Your route here

1 stop · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training · you are here ✓ understood

Picture it

Forward passRun a batch through the modelCompute lossCompare predictions to ground…Backward passBackpropagation computes gradientsUpdate weightsOptimizer steps against the…Repeat over epochs
  1. 01 Forward pass Run a batch through the model
  2. 02 Compute loss Compare predictions to ground truth
  3. 03 Backward pass Backpropagation computes gradients
  4. 04 Update weights Optimizer steps against the gradient

↺ back to 01 · Repeat over epochs

Each pass around the loop nudges the parameters to lower the loss a little; training is this loop repeated thousands of times.

Training is where a model’s behaviour comes from. A fresh neural network starts with random weights and produces nonsense. Training shows it examples, measures how wrong its outputs are, and nudges every weight slightly in the direction that would have made it less wrong. Repeat that enough times and the weights settle into values that do the task.

The measure of wrongness is the loss function. Backpropagation works out how much each weight contributed to the loss, and an optimizer based on gradient descent takes the step.

The moving parts

  • Batch: examples are processed in groups, typically tens to thousands, rather than one at a time or all at once. Each batch produces one update.
  • Epoch: one full pass through the training data. Small models often train for many epochs.
  • Learning rate: how big each step is. Too large and training diverges; too small and it crawls.
  • Checkpoints: saved copies of the weights, so a long run can resume after a crash and the best version can be kept.

Low training loss isn’t the goal

What you actually want is good performance on data the model has never seen. Driving the training loss down is only an indirect route to that, a point the standard deep learning textbook makes early in its chapter on optimization. So part of the data is held back as a validation set and checked as training runs. When training loss keeps falling but validation loss starts to rise, the model is memorizing its examples rather than learning the pattern. That’s overfitting, and the usual responses are to stop early, add regularization, or get more data.

What it costs

Training is the expensive stage of a model’s life. Every example needs a forward pass and a backward pass, and a large language model sees trillions of tokens, spread across thousands of accelerators for weeks. That makes the split between model size and data size an important budget decision. DeepMind’s 2022 Chinchilla study found that for a fixed compute budget they should grow together: double the parameters, double the training tokens. Their 70-billion-parameter model, trained on four times more data, outperformed the 280-billion-parameter Gopher on the same budget.

When training ends, the weights are frozen and the model moves on to inference: answering new inputs, with no more learning.

Where it sits

Explore nearby

In the research

All papers →

A paper that builds on Training .

Sources

  1. Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning . MIT Press, 2016; chapter 8, Optimization for Training Deep Models
  2. Hoffmann et al., "Training Compute-Optimal Large Language Models" . DeepMind, 2022 (Chinchilla)