27 ResNet
Context
Deeper ought to mean better — yet very deep networks showed rising TRAINING error (not test error, training error). That is not overfitting, it is an OPTIMIZATION problem: plain deep stacks are hard to train. He et al. (Microsoft) fix it.
The idea and the mechanism
Let a block learn not the target mapping H(x) but the RESIDUAL F(x) = H(x) − x, with the output added back to the input through an identity skip connection:
If the optimum sits close to the identity, it is easier for the network to learn F ≈ 0 than to fit the identity out of a stack of nonlinearities. That is what made 50-, 101-, 152-layer networks and deeper trainable.
calculus Why the gradient does not vanish: a highway through the skip
The derivative of a residual block's output with respect to its input:
Across L blocks the gradient is a product of such factors. Because of the I term (the identity) there is always a direct path: even if every ∂F → 0, the gradient stays ≈ 1 rather than collapsing to zero.
The "+I" is the whole point: the identity shortcut lays a highway for the gradient through hundreds of layers. Same idea as the additive path in an LSTM (ct = ct−1 + …) — there it runs through time, here through depth.
PyTorch A residual block
import torch, torch.nn as nn
class ResidualBlock(nn.Module):
def __init__(self, C):
super().__init__()
self.f = nn.Sequential(
nn.Conv2d(C, C, 3, padding=1), nn.BatchNorm2d(C), nn.ReLU(),
nn.Conv2d(C, C, 3, padding=1), nn.BatchNorm2d(C))
def forward(self, x):
return torch.relu(x + self.f(x)) # skip connection: x + F(x)
Why it matters
ResNet won ImageNet-2015 (3.57% top-5) and, at the same time, detection and localization on ImageNet plus detection and segmentation on COCO — a sweep of five competitions. Residual connections became a universal move: they sit in every Transformer block (#32), in diffusion U-Nets, nearly everywhere. "An additive path rescues the gradient" is one of the most reused principles in deep learning.
Connections
VGG pushed plain depth-stacking to its limit (19 layers) and hit degradation. ResNet removes exactly that wall — it keeps the "deeper is better" line going, but gives depth a technical way to actually work.
One trick in two disguises: an additive shortcut keeps the gradient from vanishing. In an LSTM that is ct = ct−1 + … through time; in ResNet it is y = x + F(x) through depth. The same "+1" in the derivative.
Every Transformer sublayer is wrapped in a residual plus LayerNorm. Without skip connections, training deep Transformers would be as painful as training deep CNNs was before ResNet. A 2015 idea became the standard glue of any deep architecture.
Questions worth asking
Why does simply stacking layers stop working, if it isn't overfitting?
It is an optimization problem, not a capacity one. Empirically a deep plain stack shows rising training error — it fits the data worse than a shallower one, even though it "could" just copy the shallow net and append identity layers. It turns out learning the identity out of a stack of nonlinearities is hard. The residual reformulation makes the identity the default (F = 0), and the difficulty disappears.
Is the skip connection a heuristic, or is there theory behind it?
Both. There is the "gradient highway" (see the math box). There is also another view: a ResNet behaves like an ensemble of many shallow paths of varying length (Veit, 2016) — dropping individual blocks barely breaks the network, much like dropping one tree from a forest. Yet another reading is iterative refinement of a representation. Several theories, each true in its own way.
Why addition of x + F(x) rather than concatenating features?
Addition preserves dimensionality, costs almost nothing, and yields a clean "+I" in the gradient. Concatenation (DenseNet's choice) preserves all features from all layers — richer, but it grows the channel count and the memory. It is a trade-off: ResNet is simpler and cheaper, DenseNet is more informative but heavier. Addition won on simplicity and scalability.
What to read in the original
It is a short paper — read it in full. Focus on how the "degradation problem" is framed and why the skip is about optimization rather than capacity; it is one of the most important engineering lessons in the canon.