AI History
The stories behind the ideas
Every technique we ship in production has a history — often a strange one. The math, algorithms, models, moments, people, and datasets that made modern AI.
Math
The mathematics AI is built on
Bayes' Theorem
The most important equation in machine learning was published by a dead man who never tried to publish it.
1805Least Squares
The method that trains today's neural networks was invented to find a lost dwarf planet.
1901Principal Component Analysis
In 1901, Karl Pearson figured out how to compress reality decades before there were computers to do it.
1906Markov Chains
The math behind language models was born in a feud over free will, and was first tested on a Pushkin poem.
1925P-Values
The 0.05 threshold that governs modern science was one man's rule of thumb, published in a farming manual.
1979The Bootstrap
Bradley Efron's big idea sounded like cheating: if you need more data, just resample the data you already have.
Algorithms
The methods that taught machines to learn
SHAP & Shapley Values
The math behind explaining AI predictions comes from a 1953 paper on dividing poker winnings fairly.
1957K-Means Clustering
A Bell Labs engineer designed it in 1957 to squeeze telephone signals down a wire, and it went unpublished for 25 years.
1991The Vanishing Gradient Problem
In 1991 a German student's thesis explained why deep networks refused to learn, and almost nobody read it.
1991Mixture of Experts
In 1991 researchers proposed a network that divides labor among specialists, an idea now hiding inside the biggest AI models.
1995Support Vector Machines
For a decade before deep learning, it was the sharpest tool in machine learning, and it worked by bending space.
1997No Free Lunch Theorem
A 1997 theorem proved there is no single best learning algorithm; averaged over all possible problems, every method is equally mediocre.
1998PageRank
It treated the web as a Markov chain, and became the backbone of Google.
2001Random Forests
Grow hundreds of mediocre decision trees, let them vote, and the crowd beats any single expert.
2006Monte Carlo Tree Search
A Go program that thought by playing thousands of random games to the end reinvented how machines plan.
2012Dropout
The trick that rescued deep learning came from thinking about bank tellers and sexual reproduction.
2013Word2Vec
It proved that king − man + woman ≈ queen, turning the geometry of a vector space into a map of meaning.
2014Attention
Attention was invented to help machines translate French novels, not to power chatbots.
2015Batch Normalization
It let researchers train much deeper networks much faster, and years later nobody fully agrees on why it works.
2022RLHF
The technique that made ChatGPT feel helpful learned less from data and more from human thumbs-up and thumbs-down.
Models
Architectures that changed what was possible
Hopfield Networks
In 2024, a physicist won the Nobel Prize for a neural network he had designed more than four decades earlier.
1998LeNet
By the late 1990s, a neural network called LeNet was quietly reading the handwritten digits on millions of US bank checks.
2013DQN
In 2013, a small London startup taught a single AI to play seven Atari games from nothing but the pixels on screen. By 2015, it had mastered dozens.
2014GANs
The idea for GANs reportedly came to Ian Goodfellow during a heated argument in a Montreal bar.
2017The Transformer
The architecture behind every modern chatbot was introduced in a 2017 paper cheekily titled 'Attention Is All You Need.'
2018GPT-1
Before ChatGPT had a name, there was GPT-1, a quiet 2018 experiment in reading before answering.
2018BERT
BERT learned to read by playing a fill-in-the-blank game with billions of sentences.
2020AlphaFold
In 2020, an AI solved a 50-year-old biology puzzle that had stumped scientists for generations.
2022Stable Diffusion
In the summer of 2022, image generation escaped the lab: anyone could now run a text-to-image model on their own computer.
2023LLaMA
Meta's LLaMA was meant for researchers only, but within a week its weights had leaked across the internet.
2024Reasoning Models
In 2024, AI learned a very human trick: to stop and think before answering.
2025DeepSeek-R1
In January 2025, an open model from a Chinese lab rattled markets and wiped hundreds of billions from US tech stocks in a single day.
Moments
Turning points in the story of AI
The Dartmouth Workshop
A field was founded by a summer grant proposal that promised too much and delivered a name.
1956Logic Theorist
The first AI program proved a theorem more elegantly than Bertrand Russell, and a journal refused to publish it.
1969The XOR Problem
A function a child can compute froze neural network research for over a decade.
1980Expert Systems
A program that beat doctors at diagnosing infections was never used on a single patient.
1992TD-Gammon
A neural network taught itself backgammon by playing alone, then rewrote human opening theory.
2006CUDA
NVIDIA built CUDA to sell gaming GPUs; it became the infrastructure of AI.
2016AlphaGo
Move 37: a move no human had played in 3,000 years of Go.
2020Scaling Laws
A 2020 paper found that intelligence follows a power law, and you could budget for it.
2022ChatGPT
OpenAI called it a low-key research preview; it reached 100 million users in two months.
2024MCP Protocol
Every AI app needed a custom connector for every tool, until one protocol made the math linear.
2025Vibe Coding
Karpathy named the thing everyone was already doing.
2025Agents in Production
AI stopped answering questions and started completing tasks; the demo-to-production gap was all about failure handling.
People
The minds behind the machines
Cybernetics
Norbert Wiener invented it to understand how brains, machines, and anti-aircraft guns aim at targets.
1950The Turing Test
Alan Turing sidestepped the question 'can machines think?' by proposing a parlor game instead.
1958The Perceptron
Frank Rosenblatt built a machine that learned from examples, and the Navy told reporters it would one day be conscious.
1966ELIZA
Its creator Joseph Weizenbaum was disturbed that people genuinely opened up to his chatbot.
1986Backpropagation
The idea was dismissed and rejected for years before a 1986 paper made it the engine of modern AI.
2019The Bitter Lesson
Rich Sutton condensed 70 years of AI history into a single uncomfortable claim: your clever ideas do not matter, compute does.
Datasets & Benchmarks
The data that trained — and tested — a field
Adult Income / Census
A slice of the 1994 US census became the dataset that taught machine learning about fairness.
1997California Housing
A 1997 spatial statistics paper quietly produced the dataset that would replace Boston Housing.
1998MNIST
MNIST started as a post office problem: humans sorting mail by reading zip codes by hand.
2009CIFAR-10
Hinton's students paid other students to label 60,000 tiny images, and named the result after their funder.
2011IMDB Sentiment
Stanford researchers scraped 50,000 movie reviews and made a rule: no film could appear more than 30 times.
2012ImageNet
ImageNet was built by crowdsourcing 14 million image labels at roughly a cent apiece.
2012The Titanic Dataset
Kaggle turned a 1912 shipwreck into the first machine learning problem for millions of beginners.
2015Credit Card Fraud Dataset
Two days of real European card transactions, with 492 frauds hiding among 284,807 purchases.
2015AG News
A million news articles from a forgotten academic search engine became deep learning's text classification standard.
2016SQuAD
Stanford paid crowdworkers to write 100,000 questions about Wikipedia, then watched machines pass humans within two years.
2019ARC-AGI
Francois Chollet designed puzzles a child can solve, then watched five years of AI progress barely dent them.
2020MMLU
A Berkeley PhD student assembled a 57-subject exam, from law to medicine, that became the LLM industry's report card.
2021GSM8K
OpenAI hired writers to compose 8,500 grade school word problems, and the biggest models of 2021 flunked them.
2021HumanEval
OpenAI built it to measure Codex: 164 hand-written Python problems that became the exam for every code AI since.
2023SWE-bench
Princeton asked whether AI could fix real GitHub bugs; in 2023 the best model managed under 2 percent.
2023GAIA
GAIA asked questions any careful person with a browser could answer: humans scored 92 percent, GPT-4 with plugins 15.