Era 4 · Architectures and scale · 2015

27 ResNet

Deep Residual Learning for Image Recognition · He, Zhang, Ren & Sun · Microsoft · CVPR 2016 (Best Paper)
🟥 read in full~1.5–2 horiginal ↗
The gist in 20 seconds. Very deep networks made trainable by skip connections: a block learns the residual F(x), and the output is F(x) + x. Gradient flows straight through the skip instead of vanishing, so 50–152-layer networks became trainable. Residual connections are now everywhere, the Transformer included.

Context

Deeper ought to mean better — yet very deep networks showed rising TRAINING error (not test error, training error). That is not overfitting, it is an OPTIMIZATION problem: plain deep stacks are hard to train. He et al. (Microsoft) fix it.

The idea and the mechanism

Let a block learn not the target mapping H(x) but the RESIDUAL F(x) = H(x) − x, with the output added back to the input through an identity skip connection:

y = F(x) + x

If the optimum sits close to the identity, it is easier for the network to learn F ≈ 0 than to fit the identity out of a stack of nonlinearities. That is what made 50-, 101-, 152-layer networks and deeper trainable.

calculus Why the gradient does not vanish: a highway through the skip

The derivative of a residual block's output with respect to its input:

∂y∂x = ∂(x + F(x))∂x = I + ∂F∂x

Across L blocks the gradient is a product of such factors. Because of the I term (the identity) there is always a direct path: even if every ∂F → 0, the gradient stays ≈ 1 rather than collapsing to zero.

plain net: ∂yL∂x0 = ∏l ∂Fl∂x → 0   vs   ResNet: ∏l (I + …) does not vanish

The "+I" is the whole point: the identity shortcut lays a highway for the gradient through hundreds of layers. Same idea as the additive path in an LSTM (ct = ct−1 + …) — there it runs through time, here through depth.

PyTorch A residual block
import torch, torch.nn as nn

class ResidualBlock(nn.Module):
    def __init__(self, C):
        super().__init__()
        self.f = nn.Sequential(
            nn.Conv2d(C, C, 3, padding=1), nn.BatchNorm2d(C), nn.ReLU(),
            nn.Conv2d(C, C, 3, padding=1), nn.BatchNorm2d(C))
    def forward(self, x):
        return torch.relu(x + self.f(x))   # skip connection: x + F(x)
x F(x) identity skip (a highway for the gradient) + y
The block learns the residual F(x); the input x bypasses it through the skip connection and is added to F(x). The skip is a direct highway for the gradient through depth.
Analogy. In a skyscraper you can take the stairs (through every floor, i.e. every layer) or an express elevator (the skip connection). A message (the gradient) is guaranteed to reach the ground floor from any floor even if the stairwell is blocked somewhere: the elevator always runs. And each floor only has to tweak what arrived (learn the residual) rather than rebuild the whole route.

Why it matters

ResNet won ImageNet-2015 (3.57% top-5) and, at the same time, detection and localization on ImageNet plus detection and segmentation on COCO — a sweep of five competitions. Residual connections became a universal move: they sit in every Transformer block (#32), in diffusion U-Nets, nearly everywhere. "An additive path rescues the gradient" is one of the most reused principles in deep learning.

Connections

← builds on25. VGG

VGG pushed plain depth-stacking to its limit (19 layers) and hit degradation. ResNet removes exactly that wall — it keeps the "deeper is better" line going, but gives depth a technical way to actually work.

↔ same idea, other guise11. LSTM

One trick in two disguises: an additive shortcut keeps the gradient from vanishing. In an LSTM that is ct = ct−1 + … through time; in ResNet it is y = x + F(x) through depth. The same "+1" in the derivative.

→ built into32. Transformer

Every Transformer sublayer is wrapped in a residual plus LayerNorm. Without skip connections, training deep Transformers would be as painful as training deep CNNs was before ResNet. A 2015 idea became the standard glue of any deep architecture.

Questions worth asking

Why does simply stacking layers stop working, if it isn't overfitting?

It is an optimization problem, not a capacity one. Empirically a deep plain stack shows rising training error — it fits the data worse than a shallower one, even though it "could" just copy the shallow net and append identity layers. It turns out learning the identity out of a stack of nonlinearities is hard. The residual reformulation makes the identity the default (F = 0), and the difficulty disappears.

Is the skip connection a heuristic, or is there theory behind it?

Both. There is the "gradient highway" (see the math box). There is also another view: a ResNet behaves like an ensemble of many shallow paths of varying length (Veit, 2016) — dropping individual blocks barely breaks the network, much like dropping one tree from a forest. Yet another reading is iterative refinement of a representation. Several theories, each true in its own way.

Why addition of x + F(x) rather than concatenating features?

Addition preserves dimensionality, costs almost nothing, and yields a clean "+I" in the gradient. Concatenation (DenseNet's choice) preserves all features from all layers — richer, but it grows the channel count and the memory. It is a trade-off: ResNet is simpler and cheaper, DenseNet is more informative but heavier. Addition won on simplicity and scalability.

What to read in the original

It is a short paper — read it in full. Focus on how the "degradation problem" is framed and why the skip is about optimization rather than capacity; it is one of the most important engineering lessons in the canon.