Standard 2 stops to get here · leads to 3

Subword Tokenization

Breaking words into smaller units, balancing vocabulary size with representation granularity.

Your route here

2 stops · basics first
  1. Token ✓ understood

    The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

  2. Tokenization ✓ understood

    Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.

  3. Subword Tokenization · you are here ✓ understood

Breaking words into smaller units, balancing vocabulary size with representation granularity.

This concept is essential for understanding large language models and forms a key part of modern AI systems.

  • Tokenization
  • BPE
  • WordPiece

Where it sits

Before this

Tokenization
Subword Tokenization

Explore nearby