Skip to content
StrataHub

Datasets & Benchmarks · 2012

ImageNet

ImageNet was built by crowdsourcing 14 million image labels at roughly a cent apiece.

In 2006, Fei-Fei Li, then a young professor, made a contrarian bet: computer vision was stuck not because algorithms were weak but because datasets were tiny. She set out to label an image for every noun in WordNet, the linguistic database of English concepts. Senior colleagues warned the project was too big and would stall her career.

Her students estimated that labeling the images by hand would take undergraduates about 19 years. The solution came from an unexpected place: Amazon Mechanical Turk, where nearly 50,000 workers across 167 countries labeled and verified images for pennies per judgment. By 2009, when ImageNet was presented as a poster at the CVPR conference, it already held around 3.2 million labeled images across some 5,000 categories, and it would eventually grow past 14 million images in over 20,000 categories.

From 2010 the team ran an annual competition, the ImageNet Large Scale Visual Recognition Challenge, on a 1,000-category subset. For two years progress was incremental, with error rates hovering above 25 percent.

Then came 2012. Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton entered AlexNet, a deep convolutional neural network trained on two consumer gaming GPUs, using dropout to fight overfitting. It scored 15.3 percent top-5 error against 26.2 percent for the runner-up, a gap so large it forced the entire field to change direction almost overnight.

That result is widely treated as the starting gun of the deep learning era. Within five years every winning entry was a deep network, error rates dropped below 3 percent, past typical human performance, and the challenge retired in 2017. ImageNet proved a thesis the field has kept relearning ever since: scale of data can matter as much as cleverness of algorithm.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.