Landmark 5 stops to get here · leads to 7

Pre-training

Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.

Your route here

5 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  4. Unsupervised Learning ✓ understood

    Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.

  5. Self-Supervised Learning ✓ understood

    Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).

  6. Pre-training · you are here ✓ understood

Picture it

  1. 01 Huge unlabeled dataset e.g. trillions of tokens of web text
  2. 02 Self-supervised objective Predict the next or the masked token
  3. 03 Base model General knowledge, no task focus yet
  4. 04 Fine-tuning Small labeled or task-specific dataset
  5. 05 Specialized model e.g. a chat assistant or classifier
Notice where the cost goes: pre-training does the expensive general learning once, so each fine-tune afterwards is cheap.

Where it sits

Explore nearby

In the research

All papers →

4 papers that build on Pre-training .