16 AlexNet
Context
By 2012 everything was ready separately: the ImageNet dataset (#15), the CNN idea (#9), backprop. What was missing was compute and a couple of tricks. Krizhevsky, Sutskever and Hinton put it all together.
The idea and the mechanism
A deep CNN: 5 convolutional + 3 fully connected layers, ~60M parameters. What made it work: ReLU (does not saturate → training several times faster than with tanh), training on 2 GPUs (the model was split across the cards — otherwise it did not fit into 3 GB of memory), dropout in the fully connected layers, data augmentation (crops, reflections, color shifts), overlapping max-pooling and local response normalization.
calculus Why ReLU beat the sigmoid: the vanishing gradient
The derivative of the sigmoid is bounded and it saturates:
In a deep network backprop multiplies these derivatives layer after layer, so the gradient shrinks exponentially:
For ReLU f(z) = max(0, z) the derivative is 1 on the active part (z > 0) and 0 on the inactive one. On active neurons the gradient does not shrink — there is no multiplicative decay through depth. That is what made it possible to train genuinely deep networks. The price is "dead" neurons (if z < 0 always, the gradient is 0 forever), later treated with variants such as Leaky ReLU.
PyTorch An AlexNet-style block
import torch.nn as nn
block = nn.Sequential(
nn.Conv2d(3, 96, kernel_size=11, stride=4), # large filters at the input
nn.ReLU(inplace=True), # faster than tanh, does not saturate
nn.MaxPool2d(kernel_size=3, stride=2), # overlapping pooling
)
# the full network: 5 conv blocks + 3 fc, dropout in the fc part, ~60M parameters
Why it matters
The moment deep learning took off. It proved empirically that deep CNNs + GPUs + large data give accuracy that hand-crafted features cannot reach. After 2012 the whole of computer vision (and soon the rest of ML) switched to deep networks, and demand for GPUs began — which is what gave NVIDIA its present role.
Connections
Architecturally AlexNet is LeNet grown deeper: the same convolutions + pooling, but more layers, ReLU instead of the sigmoid and dropout. The 1998 idea waited for data and hardware — and fired without any change to its essence.
Without ImageNet's scale a deep network would have overfitted or shown no advantage. The dataset is the fuel without which the engine would not have moved; AlexNet's win vindicated Fei-Fei Li's bet on data in retrospect.
AlexNet opened the "deeper = better" race. But simply stacking layers stopped working (degradation). Three years later ResNet solves that with skip connections and pushes depth to 152 layers — a continuation of the trajectory that starts here.
Questions worth asking
What was decisive — ReLU, GPUs, dropout or the data?
All of them are necessary, but their roles differ. GPUs + ImageNet made depth trainable at scale; ReLU made training fast; dropout and augmentation kept overfitting in check. Remove any one and the result falls apart. The deepest enabler is the meeting of cheap parallel compute (GPU/CUDA) with large data; the tricks only made that meeting productive.
Why 2012, if CNNs had existed since 1998?
The idea was waiting for its substrate. In 1998 there was no dataset of that scale, no GPU compute, no ReLU/dropout; deep CNNs on weak hardware and small data lost to SVMs. ImageNet and CUDA cards lifted both constraints almost simultaneously — and the accumulated idea cashed out instantly. A typical DL storyline: the right thought waits for data and compute.
Splitting across 2 GPUs — a fundamental decision or a workaround?
A workaround for the memory limit (3 GB per card) — and an early example of model parallelism. Curiously, the forced weak coupling between the two "halves" on separate GPUs had a mild regularizing effect. Today a single AlexNet needs none of this, but the same move (split the model across devices) is now mandatory for training giant LLMs.
What to read in the original
Short (9 pages) and historical — worth reading in full. Pay attention to the engineering details (memory, 2 GPUs, ReLU): an excellent example of how system constraints shape an architecture.