Training is where a model’s behaviour comes from. A fresh neural network starts with random weights and produces nonsense. Training shows it examples, measures how wrong its outputs are, and nudges every weight slightly in the direction that would have made it less wrong. Repeat that enough times and the weights settle into values that do the task.
The measure of wrongness is the loss function. Backpropagation works out how much each weight contributed to the loss, and an optimizer based on gradient descent takes the step.
The moving parts
- Batch: examples are processed in groups, typically tens to thousands, rather than one at a time or all at once. Each batch produces one update.
- Epoch: one full pass through the training data. Small models often train for many epochs.
- Learning rate: how big each step is. Too large and training diverges; too small and it crawls.
- Checkpoints: saved copies of the weights, so a long run can resume after a crash and the best version can be kept.
Low training loss isn’t the goal
What you actually want is good performance on data the model has never seen. Driving the training loss down is only an indirect route to that, a point the standard deep learning textbook makes early in its chapter on optimization. So part of the data is held back as a validation set and checked as training runs. When training loss keeps falling but validation loss starts to rise, the model is memorizing its examples rather than learning the pattern. That’s overfitting, and the usual responses are to stop early, add regularization, or get more data.
What it costs
Training is the expensive stage of a model’s life. Every example needs a forward pass and a backward pass, and a large language model sees trillions of tokens, spread across thousands of accelerators for weeks. That makes the split between model size and data size an important budget decision. DeepMind’s 2022 Chinchilla study found that for a fixed compute budget they should grow together: double the parameters, double the training tokens. Their 70-billion-parameter model, trained on four times more data, outperformed the 280-billion-parameter Gopher on the same budget.
When training ends, the weights are frozen and the model moves on to inference: answering new inputs, with no more learning.