Moments · 1992
TD-Gammon
A neural network taught itself backgammon by playing alone, then rewrote human opening theory.
In 1992, IBM researcher Gerald Tesauro unveiled TD-Gammon, a backgammon program that learned almost entirely by playing against itself. It started with essentially random play and improved through hundreds of thousands of self-play games, with no database of expert moves to imitate.
The engine was a neural network trained with temporal-difference learning, a reinforcement learning method developed by Richard Sutton. After each move, the network nudged its evaluation of the previous position toward its evaluation of the new one, so that the eventual win or loss propagated backward through the whole game. Good positions were discovered, not taught.
The result stunned both AI researchers and backgammon professionals. TD-Gammon reached a level close to the best human players in the world, and its judgment in certain positions was strong enough that top players began studying it.
Most remarkably, it changed how humans play. In some standard opening situations, TD-Gammon preferred plays that conventional wisdom rejected, such as splitting the back checkers where experts had favored the aggressive 'slotting' play. Analysis vindicated the program, and expert practice shifted, an early case of humans learning strategy from a machine.
For years TD-Gammon was a puzzling one-off: the same techniques failed to conquer chess or Go, and researchers debated whether backgammon's dice rolls made it uniquely suited to self-play. Two decades later the approach returned at scale. DeepMind's DQN and AlphaGo both descend directly from TD-Gammon's recipe of neural network evaluation plus reinforcement learning through self-play.
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.