The number of distinct tokens a language model can process, typically 30K-100K+ tokens.
This concept is essential for understanding large language models and forms a key part of modern AI systems.
Related Concepts
- Tokenization
- Token
- Embedding
The number of distinct tokens a language model can process, typically 30K-100K+ tokens.
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
The number of distinct tokens a language model can process, typically 30K-100K+ tokens.
This concept is essential for understanding large language models and forms a key part of modern AI systems.