Era 6 · Generative models and systems · 2020

46 DDPM (Diffusion)

Denoising Diffusion Probabilistic Models · Ho, Jain & Abbeel · UC Berkeley · NeurIPS
🟥 read in full~2–3 horiginal ↗
The gist in 20 seconds. Train a network to "denoise" data step by step: a forward process destroys the image with noise, the reverse process undoes it. The loss collapses to a plain MSE between the added noise and the predicted noise. GAN-level quality without GAN instability. The start of the diffusion era.

Context

GANs (#20) gave sharp images but were temperamental (instability, mode collapse). Ho, Jain and Abbeel make diffusion practical and good.

The idea and the mechanism

Forward (fixed, nothing learned): over T steps, gradually add Gaussian noise until only pure noise is left. Reverse (learned): the network learns to UNDO that step by step. Generation: start from pure noise and denoise iteratively into an image.

probability From the variational bound to a plain MSE on the noise

Forward. Adding noise has a closed form for any step (a property of Gaussians): you can get xt from x0 in one go:

xt = √ ᾱt  · x0 + √ 1−ᾱt  · ε,   ε ∼ N(0, I)

where ᾱt = ∏s≤t(1−βs). Reverse. We learn pθ(xt−1 | xt). The full variational lower bound (ELBO) is unwieldy, but the authors show that it reduces to a surprisingly simple form — predict the added noise:

Lsimple = Et, x0, ε ‖ ε − εθ(xt, t) ‖²

That is, the network εθ looks at the noised xt and the step t and predicts which noise was mixed in — training is an ordinary MSE. The beauty of it is that a hard generative problem has turned into "guess the noise".

PyTorch A diffusion training step
import torch

def diffusion_loss(x0, model, abar, T):
    t   = torch.randint(0, T, (len(x0),))
    eps = torch.randn_like(x0)
    a   = abar[t].view(-1, 1, 1, 1)
    xt  = a.sqrt() * x0 + (1 - a).sqrt() * eps   # add the noise in one step
    return ((eps - model(xt, t)) ** 2).mean()    # predict the noise (MSE)
x₀ xₜ noise forward: add noise → ← reverse: the network denoises (learned) generation:from noise → image
The forward process destroys the image into noise by a fixed rule; the learned reverse process walks it back. To generate = start from noise and denoise.
Analogy. Imagine watching, over and over, a sharp photograph dissolve into television snow. If you learn to undo each small step of that decay — to take off just a little noise — then, starting from pure snow, you can bring a meaningful picture out of it step by step. Diffusion learns exactly this micro-denoising.

Why it matters

It gave GAN-quality generation WITHOUT the instability and with full mode coverage (diversity) → diffusion became the dominant paradigm for image generation (and audio, video, molecules). The direct foundation of Stable Diffusion (#47), DALL·E 2, Imagen.

Connections

↔ successor20. GAN

The same task (image generation), the opposite approach: instead of two networks competing, one network learns to denoise. Diffusion cured the GAN's signature illnesses (instability, mode collapse) and pushed it out of image generation by the early 2020s.

DDPM works in pixels — expensive. Stable Diffusion moves the same math into a compact latent space, putting text-to-image within reach of consumer GPUs. A direct practical continuation.

↔ echoes30. WaveNet

Both generate iteratively, but differently: WaveNet is autoregressive, one sample at a time; diffusion works on the whole image in parallel, but over many denoising steps. Diffusion partly cures the slowness of sequential generation (though it still needs dozens of steps of its own).

Questions worth asking

Why predict the noise rather than the clean image directly?

Mathematically it is equivalent (given the noise and xt you can recover x0), but numerically ε-prediction gives a better-conditioned problem: the target has unit variance at every step, so the gradients are steadier. Empirically, predicting the noise trains noticeably better than predicting the image — a lucky choice of parameterization.

Diffusion needs dozens or hundreds of denoising steps — does that not kill the speed?

It does, and that is the main price paid against a GAN (a single pass). Hence the accelerators: DDIM (deterministic sampling in fewer steps), distillation into few-step or one-step models, consistency models. So "many steps" is a real weakness, actively being worked on; the quality and robustness of diffusion outweighed the slowness.

Where did the idea come from — was it invented from scratch in 2020?

No: the theoretical basis is non-equilibrium thermodynamics (Sohl-Dickstein, 2015), and the roots run back into score matching and stochastic differential equations. DDPM's contribution was making the approach practical: simplifying it down to ε-prediction and an MSE loss turned an elegant but unwieldy theory into a working recipe at SOTA quality.

What to read in the original

Read it in full — the math (the forward process as a marginal, the simplification to ε-prediction) is both important and lovely; it is the best way to see why "train a denoiser" = "train a generative model".