Reference 3 stops to get here

BPE

Byte Pair Encoding - a subword tokenization algorithm that iteratively merges frequent character pairs to create a vocabulary.

Your route here

3 stops · basics first
  1. Token ✓ understood

    The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

  2. Tokenization ✓ understood

    Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.

  3. Subword Tokenization ✓ understood

    Breaking words into smaller units, balancing vocabulary size with representation granularity.

  4. BPE · you are here ✓ understood

Where it sits

BPE

Leads to

Nothing yet: a destination in its own right.

Explore nearby