Standard 1 stop to get here · leads to 1
Benchmark
A standardized dataset and task used to compare model performance across different approaches (ImageNet, GLUE, SuperGLUE).
Your route here
1 stop · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Benchmark · you are here ✓ understood
Where it sits
Explore nearby
Evaluation Baseline Model A simple reference model (random, majority class, simple heuristic) used to benchmark more complex models against. Evaluation Benchmark Gaming Optimizing models specifically for benchmark performance rather than real-world capabilities, inflating scores artificially. Vision & Multimodal ImageNet A large-scale dataset of 14M images in 20K categories, historically used as the benchmark for image classification models. Evaluation Test Set A final portion of data unseen during training and validation, used for unbiased evaluation of model performance. Language & LLMs Emergent Abilities Capabilities that appear suddenly in large language models at certain scales, not present in smaller models.
In the research
All papers →7 papers that build on Benchmark ; showing 5, canon first.