Reference 1 stop to get here · leads to 2
N-gram
A contiguous sequence of n items (words, characters) from text, used in language modeling and feature extraction.
Your route here
1 stop · basics first
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- N-gram · you are here ✓ understood
Where it sits
Explore nearby
Language & LLMs Language Modeling Learning probability distributions over sequences of words to predict what comes next. Language & LLMs TF-IDF Term Frequency-Inverse Document Frequency - a statistical measure of word importance in documents, used for information retrieval. Evaluation Perplexity A metric measuring how well a language model predicts text - lower perplexity indicates better prediction. Language & LLMs Tokenization Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it. Foundations Hidden Markov Model A statistical model with hidden states that transition probabilistically, generating observable outputs.
In the research
All papers →2 papers that build on N-gram .