ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng et al. · CVPR 2009
doi:10.1109/CVPR.2009.5206848
In short
The authors build a database of millions of labelled images organised by the WordNet hierarchy, labelled at scale through crowdsourcing. It is far larger and more varied than earlier vision datasets.
Why it matters
ImageNet and its annual challenge gave deep learning the benchmark it needed to prove itself.
Read first
The 4 Field Guide ideas this paper leans on.
Starting from scratch? The full route 10 ideas · basics first
- Dataset · read first ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Supervised Learning ✓ understood
Learning from examples paired with the correct answer, so a model can predict answers for new inputs it hasn't seen.
- Classification ✓ understood
A supervised learning task where the model assigns each input to one of a fixed set of categories, such as spam or not spam.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Deep Learning ✓ understood
A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.
- Computer Vision ✓ understood
The field of AI that gets computers to extract meaning from images and video: what is in them, where it is, and how it moves.
- Image Classification · read first ✓ understood
Assigning a single label or category to an entire image, a fundamental computer vision task.
- Benchmark · read first ✓ understood
A standardized dataset and task used to compare model performance across different approaches (ImageNet, GLUE, SuperGLUE).
- ImageNet · read first ✓ understood
A large-scale dataset of 14M images in 20K categories, historically used as the benchmark for image classification models.