BLEU: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni et al. · ACL 2002
doi:10.3115/1073083.1073135
In short
BLEU scores a machine translation by how many of its word sequences also appear in human reference translations, with a penalty for being too short. It is quick, cheap, language-independent and tracks human judgements well enough to compare systems.
Why it matters
It showed that an automatic metric can drive a field’s progress, and it is still reported in translation papers.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 8 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Recurrent Neural Network ✓ understood
A neural network architecture with loops that allow information to persist, designed for sequential data like text and time series.
- Sequence-to-Sequence ✓ understood
Models that transform input sequences to output sequences, used for translation, summarization, and generation.
- Machine Translation · read first ✓ understood
Automatically translating text from one language to another using neural models (typically encoder-decoder architectures).
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- N-gram · read first ✓ understood
A contiguous sequence of n items (words, characters) from text, used in language modeling and feature extraction.
- BLEU Score · read first ✓ understood
A metric for evaluating machine translation quality by comparing n-gram overlap between generated and reference text.