Reference 2 stops to get here

Vocabulary Size

The number of distinct tokens a language model can process, typically 30K-100K+ tokens.

Your route here

2 stops · basics first
  1. Token ✓ understood

    The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

  2. Tokenization ✓ understood

    Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.

  3. Vocabulary Size · you are here ✓ understood

The number of distinct tokens a language model can process, typically 30K-100K+ tokens.

This concept is essential for understanding large language models and forms a key part of modern AI systems.

  • Tokenization
  • Token
  • Embedding

Where it sits

Before this

Tokenization
Vocabulary Size

Leads to

Nothing yet: a destination in its own right.

Explore nearby