Era 3 · The deep learning explosion · 2014

19 Dropout

Dropout: A Simple Way to Prevent Neural Networks from Overfitting · Srivastava, Hinton, Krizhevsky, Sutskever, Salakhutdinov · JMLR
🟧 read selectively~40 minoriginal ↗
The gist in 20 seconds. At every training step we randomly "switch off" neurons → the network stops relying on their joint conspiracies and learns robust features. Equivalent to averaging an exponential ensemble of thinned subnetworks. A cheap, powerful regularizer that became the standard.

Context

Large networks overfit — they memorize noise and generalize badly. Srivastava, Hinton et al. offer a simple, general regularizer.

The idea and the mechanism

At every training step each neuron is switched off (zeroed) with probability 1−p. The network cannot rely on particular neurons and their joint co-adaptation → it is forced to learn redundant, robust features. At inference all neurons are active and the outputs are rescaled so the expected magnitude is preserved.

probability Dropout as ensemble averaging

Take a neuron with activation a that is kept with probability p: the mask is m ∼ Bernoulli(p), the output y = m · a. Then the expected output is:

E[y] = p · a

So at test time, when every neuron is on, the output is multiplied by p — to make the average magnitude match training (or you use "inverted dropout": divide by p during training and leave test time alone).

The main interpretation. Each step trains one of 2n "thinned" subnetworks (one per subset of the n neurons), and they all share weights. Test-time rescaling approximately computes the geometric mean of the predictions of that entire exponential ensemble. In other words, dropout ≈ training a giant ensemble at the cost of a single network.

NumPy Inverted dropout (train / eval)
import numpy as np

def dropout(a, p=0.5, train=True):
    if not train:
        return a                              # at inference — unchanged
    mask = (np.random.rand(*a.shape) < p) / p # mask + 1/p scaling
    return a * mask                           # some neurons are zeroed
at each step some neurons are randomly switched off
The crossed-out neurons are off at this step. Every step gets its own random "thinned" subnetwork; all of them share the same weights.
Analogy. Hinton came up with dropout in a bank: he noticed the tellers were constantly being rotated, and was told why — to make it harder for them to collude on fraud. If any employee can be pulled at any moment, nobody can build a fragile scheme that depends on specific people. Same with neurons: since you might be "switched off", you cannot lean on your neighbour — you have to be useful yourself.

Why it matters

Cheap, general, and for a long time the standard in fully connected and recurrent networks; conceptually it ties regularization to ensembling. In modern CNNs it has partly been displaced by BatchNorm and by sheer volumes of data, but in Transformers (dropout in attention/FFN) it lives on.

Connections

← applied in16. AlexNet

Dropout in the fully connected layers was one of the key tricks that kept the 60M-parameter AlexNet from overfitting ImageNet. The regularizer and the 2012 breakthrough came out of the same lab and worked together.

Both stabilize/regularize training, but differently: dropout injects noise by switching neurons off, BatchNorm normalizes activations (and throws in mild regularization as a side effect). With the arrival of BatchNorm and large datasets, dropout was pushed well back in convolutional networks.

↔ same idea, other guise12. Random Forests

In both cases the strength comes from an ensemble of diverse models. A forest builds many separate trees; dropout hides an exponential ensemble of subnetworks inside a single network with shared weights. Different routes to the same wisdom of crowds.

Questions worth asking

Why does test-time scaling by p only approximate the ensemble rather than compute it exactly?

Exact averaging would mean running all 2n subnetworks and averaging them — impossible. Multiplying by p is a careful approximation to the geometric mean of their outputs, exact for linear networks and good for moderate nonlinearities. In practice the approximation works very well, even though it is not strictly equal to the ensemble.

Why has dropout all but disappeared from modern CNNs, yet stayed in Transformers?

In CNNs its role as a regularizer was largely taken over by BatchNorm, augmentation and abundant data, and dropout between convolutional maps interfered with BN's statistics. Transformers have no BatchNorm (they use LayerNorm), the models are huge and prone to overfitting particular data — so dropout in attention and the FFN remains useful. The choice of regularization depends on the architecture.

Does switching neurons off at random every step hurt convergence?

It slows it down — training is noisier and needs more epochs (you are effectively training many subnetworks at once). It is a deliberate trade: slightly slower training for markedly better generalization. So dropout is used where overfitting is a real threat (big networks, little data), and dropped when data is plentiful.

What to read in the original

The write-up plus the key passages is enough: the masking idea, the test-time scaling and the ensemble interpretation. The long empirical section can be skimmed.