Standard 1 stop to get here · leads to 1

Benchmark

A standardized dataset and task used to compare model performance across different approaches (ImageNet, GLUE, SuperGLUE).

Your route here

1 stop · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Benchmark · you are here ✓ understood

Where it sits

Before this

Dataset
Benchmark

Explore nearby

In the research

All papers →

7 papers that build on Benchmark ; showing 5, canon first.

Canon · 2009 ImageNet: A Large-Scale Hierarchical Image Database ImageNet and its annual challenge gave deep learning the benchmark it needed to prove itself. Canon · 2020 Measuring Massive Multitask Language Understanding It became the headline benchmark in LLM release notes for years, and a case study in benchmarks saturating. Frontier · Sep 2025 · 282 citations Why Language Models Hallucinate It reframes hallucination as an incentive problem baked into how we measure models, not a mystery. Frontier · Feb 2026 · 174 citations SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks It is the first careful evidence that skills work, and that smaller models with good skills can match bigger ones. Frontier · Oct 2025 · 108 citations GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks It measures economic usefulness directly, instead of exam-style puzzles.