Algorithms · 1991
The Vanishing Gradient Problem
In 1991 a German student's thesis explained why deep networks refused to learn, and almost nobody read it.
Neural networks learn by propagating an error signal backward from the output, adjusting each weight in proportion to its gradient. In deep networks, that signal has to pass through many layers, and something went quietly wrong along the way.
In 1991, Sepp Hochreiter, a student of Jurgen Schmidhuber in Munich, analyzed this in his diploma thesis. He showed that as gradients are multiplied layer by layer, they tend to shrink toward zero, or occasionally blow up. The early layers, farthest from the error, receive almost no usable signal and effectively stop learning.
The problem was worst in recurrent networks trying to connect events far apart in time; the influence of an input decayed exponentially the longer the network had to remember it. Yoshua Bengio and colleagues documented the same difficulty in 1994.
For years this was a major reason deep and recurrent networks had a reputation for being nearly untrainable. The vanishing gradient was the invisible wall.
The fixes came gradually: the LSTM that Hochreiter and Schmidhuber designed in 1997 to preserve signals over time, better activation functions like ReLU, careful initialization, and later batch normalization and residual connections.
Understanding the vanishing gradient was a prerequisite for the deep learning revolution. You cannot fix a wall you cannot see, and Hochreiter's overlooked thesis is where the field first saw it clearly.
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.