14 Deep Belief Nets
Context
2006: deep networks are formally trainable by backprop, but in practice they train badly — vanishing gradients, poor minima, not enough data; they are considered impractical. Hinton and his co-authors offer a way around it and revive the field's interest.
The idea and the mechanism
A deep generative network built as a stack of RBMs (restricted Boltzmann machines). Train greedily, layer by layer and without supervision: the first RBM models the data; its hidden activations become the "data" for the second RBM; and so on. Then the whole network is carefully fine-tuned (supervised backprop, for instance). Layer-wise pre-training gives the weights a good initialization, from which fine-tuning does converge.
probability RBMs and the contrastive divergence trick
An RBM defines a joint distribution over visible units v and hidden units h through an energy:
We want to maximize the likelihood of the data. The gradient of the log-likelihood has an elegant but intractable form — the difference of two expectations:
The second expectation requires sampling from the model itself (expensive MCMC). Contrastive Divergence approximates it with a single Gibbs step starting from the data (a "reconstruction") — and that is what made training fast. The RBMs are then stacked greedily: each layer learns the distribution of the previous layer's activations.
NumPy One step of contrastive divergence (CD-1)
import numpy as np
sig = lambda z: 1 / (1 + np.exp(-z))
def cd1(v, W, eta=0.1): # training a single RBM
h = (sig(v @ W) > np.random.rand(W.shape[1])).astype(float) # data
v2 = sig(h @ W.T) # reconstruction of the visible layer
h2 = sig(v2 @ W)
W += eta * (np.outer(v, h) - np.outer(v2, h2)) # data − reconstruction
return W
Why it matters
The paper credited with starting the modern "deep learning" era: it showed that deep networks can be trained and that they give SOTA results (on MNIST), which handed them back their respectability. The Hinton–Bengio–LeCun group held together around it (on a CIFAR grant) through the lean years. A caveat: the method itself (RBMs + layer-wise pre-training) was soon superseded — ReLU, better initialization, BatchNorm and an abundance of data made it possible to train deep networks with plain backprop.
Connections
The RBM is a relative of the Hopfield network and the Boltzmann machine: the same energy-based models out of statistical physics, but stochastic and trainable. The "energy → probability → learning" line runs straight from here.
DBNs kindled the belief that deep networks are trainable and pulled a community together. Six years later AlexNet proved it without any RBMs — straight backprop on a GPU. Pre-training turned out to be a crutch that stopped being needed, but it is what carried the field through the dark period.
The DBN answered "how do you get training of a deep network off the ground at all?" with pre-training. BatchNorm (along with good initialization and ReLU) solved the same problem differently — by stabilizing the training process itself, which made layer-wise pre-training unnecessary.
Questions worth asking
If pre-training helped so much, why was it abandoned?
Because the reason it was needed — deep networks training badly — was removed. ReLU (no saturation), sensible initialization (Xavier/He), BatchNorm and large datasets let the gradient flow through depth directly. Once backprop worked head-on, the detour was redundant.
How is RBM pre-training different from autoencoder pre-training?
Both are unsupervised initialization through reconstruction, and the "a layer learns a representation of its input" idea is shared. An RBM is a probabilistic energy-based model trained by contrastive divergence; an autoencoder is a deterministic network minimizing reconstruction error directly with backprop. Autoencoders are simpler and quickly displaced RBMs as the way to pre-train — until they too became unnecessary.
Did this paper "coin" the term deep learning?
No — the term is older. It popularized the modern framing and became the symbolic start of the era, but "deep learning" was in use before that. Crediting it with coining the term is a common mistake; "the paper that brought depth back into the mainstream" is the accurate description.
What to read in the original
This write-up is enough — the historical role matters more than the mechanics of RBMs, which are niche today. If you want to dig, look at the idea of greedy layer-wise training and at contrastive divergence (the math box).