Algorithms · 2012
Dropout
The trick that rescued deep learning came from thinking about bank tellers and sexual reproduction.
Around 2012, Geoffrey Hinton and his students faced a stubborn problem: large neural networks memorized their training data instead of learning to generalize. The fix they proposed, called dropout, was almost absurdly simple.
During each training step, randomly switch off a fraction of the neurons, say half, so the network can never rely on any single unit being present. Hinton later said the idea came partly from a bank where the tellers kept rotating, which he suspected was to stop any small group from colluding, and partly from why sexual reproduction shuffles genes so no gene depends too much on a fixed set of partners.
Forcing neurons to work with random subsets of their neighbors prevents them from co-adapting into fragile, over-specialized teams. Each unit has to learn a feature that is useful on its own, which makes the whole network more robust.
There is a second way to see it. Dropout effectively trains an enormous ensemble of thinned networks that share weights, then averages them at test time, echoing the same wisdom-of-crowds effect behind random forests.
The 2012 paper by Hinton, Srivastava, Krizhevsky, Sutskever, and Salakhutdinov, and the fuller 2014 journal version, made dropout a standard ingredient. It was part of the toolkit that let networks like AlexNet win on ImageNet.
Newer architectures lean on other forms of regularization, but dropout remains one of the most elegant examples of how a small, cheap idea can unlock much larger models.
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.