A large language model is a neural network trained to do one thing: given some text, predict the next token. Do that across a large share of the web, books and code, and the model picks up grammar, facts, styles of argument and how programs fit together, because all of them help it guess what comes next.
Generating text is that same prediction run in a loop. The model scores every token in its vocabulary, one is picked, it’s appended to the input, and the whole network runs again. A 500-word answer is several hundred of these passes.
Nearly every modern LLM is a transformer, the architecture introduced in 2017, whose attention mechanism lets each token draw on every other token in the input.
How one gets built
- Pre-training: the model learns next-token prediction on enormous amounts of raw text. The text supplies its own labels, so no one has to annotate it. This is where nearly all the compute goes.
- Fine-tuning: a much smaller round on curated examples teaches it to follow instructions and hold a conversation.
- Learning from preferences: people rank the model’s answers and the model is tuned toward the ones they prefer, most famously with RLHF. In OpenAI’s 2022 InstructGPT paper, labelers preferred a 1.3-billion-parameter model tuned this way over the 175-billion-parameter GPT-3.
Why scale mattered
The 2020 GPT-3 paper showed that a big enough model could take on a new task from a few examples written into the prompt, with no retraining at all. That is in-context learning, and the authors found it improved sharply as models grew. It’s why one model can translate, summarize, classify and write code: the task is described in the prompt rather than trained in.
The catch
- It predicts plausible text, not checked text. When it doesn’t know, it can produce a fluent, confident, wrong answer: a hallucination.
- It only sees its context window. Everything it uses must fit in that window, and every token is paid for on every call. The Duel makes that cost concrete.
- Its knowledge stops at its training cutoff, unless you hand it documents or tools at run time.