Skip to content
StrataHub

Datasets & Benchmarks · 1998

MNIST

MNIST started as a post office problem: humans sorting mail by reading zip codes by hand.

In the late 1980s the US Postal Service faced an expensive bottleneck. Millions of hand-addressed envelopes had to be sorted every day, and no machine could reliably read handwriting. At AT&T Bell Labs, Yann LeCun and his colleagues built neural networks to read handwritten zip code digits, and by the mid-1990s descendants of that work were reading a meaningful share of the checks deposited in the United States.

Progress needed a fair, shared benchmark. LeCun, Corinna Cortes and Christopher Burges assembled MNIST in 1998 from two handwriting collections gathered by NIST, the US standards agency: one written by Census Bureau employees, the other by high school students.

The M stands for modified, and the modification mattered. NIST's original split trained on tidy digits from professional census workers and tested on messier student handwriting, which made every model look worse than it was. The MNIST authors remixed the writers into a 60,000-image training set and a 10,000-image test set, with each digit centered and scaled into a 28 by 28 pixel grayscale square.

The dataset debuted alongside LeNet-5 in the 1998 paper Gradient-Based Learning Applied to Document Recognition, one of the founding documents of convolutional neural networks. For the next two decades MNIST was the fruit fly of machine learning: small enough to run anywhere, standardized enough that a single error percentage let researchers compare methods across decades.

Error rates fell from around 1 percent with early convolutional networks to under 0.25 percent with modern methods, close to the limit set by genuinely ambiguous digits. Today MNIST is considered too easy to prove anything, yet it remains the hello world of machine learning: the first dataset most practitioners ever train a model on.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.