Evaluation Jul 2002

BLEU: a Method for Automatic Evaluation of Machine Translation

Kishore Papineni et al. · ACL 2002

doi:10.3115/1073083.1073135

In short

BLEU scores a machine translation by how many of its word sequences also appear in human reference translations, with a penalty for being too short. It is quick, cheap, language-independent and tracks human judgements well enough to compare systems.

Why it matters

It showed that an automatic metric can drive a field’s progress, and it is still reported in translation papers.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 8 ideas · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Recurrent Neural Network ✓ understood

    A neural network architecture with loops that allow information to persist, designed for sequential data like text and time series.

  4. Sequence-to-Sequence ✓ understood

    Models that transform input sequences to output sequences, used for translation, summarization, and generation.

  5. Machine Translation · read first ✓ understood

    Automatically translating text from one language to another using neural models (typically encoder-decoder architectures).

  6. Token ✓ understood

    The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

  7. N-gram · read first ✓ understood

    A contiguous sequence of n items (words, characters) from text, used in language modeling and feature extraction.

  8. BLEU Score · read first ✓ understood

    A metric for evaluating machine translation quality by comparing n-gram overlap between generated and reference text.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers