Era 6 · Generative models and systems · 2022

47 Stable Diffusion (LDM)

High-Resolution Image Synthesis with Latent Diffusion Models · Rombach, Blattmann et al. · LMU/Runway · CVPR
🟧 read selectively~1.5–2 horiginal ↗
The gist in 20 seconds. Moves diffusion into the latent space of a pre-trained autoencoder and adds conditioning through cross-attention, cutting the compute by orders of magnitude. The architecture that became Stable Diffusion — text-to-image on consumer GPUs.

Context

DDPM (#46) works in PIXEL space → expensive: for 512×512, each of hundreds of denoising steps runs at full resolution. Rombach et al. make diffusion cheap.

The idea and the mechanism

Move diffusion into a COMPACT latent space. (1) Train a perceptual autoencoder that compresses an image into a small latent and back. (2) Run the diffusion in that latent — orders of magnitude less compute. (3) The decoder restores the pixels. Conditioning (text, a mask) is fed in through cross-attention inside the denoiser's U-Net.

probability Latent diffusion + cross-attention conditioning

Two stages. First the autoencoder: z = E(x) compresses a 512×512×3 image into a latent of, say, 64×64×4 — about 48 times fewer elements. The diffusion (the same ε-prediction as in DDPM) runs in z rather than in pixels:

L = E ‖ ε − εθ(zt, t, c) ‖²,   then x = D(z0)

Conditioning. The text prompt is encoded into embeddings c, which enter the denoiser through cross-attention: the latent features supply the Query, the text supplies Key/Value:

Attention(Qlatent, Ktext, Vtext)

So every region of the image being generated "looks at" the relevant words of the prompt. Moving into the latent is where the main win comes from (orders of magnitude cheaper); cross-attention is what makes it steerable by text.

PyTorch Latent diffusion (outline)
z0 = encoder(x)                      # compress the image into a latent
zt = a.sqrt()*z0 + (1-a).sqrt()*eps  # add noise in the latent
pred = unet(zt, t, cond=text_emb)    # denoiser + cross-attention to the text
loss = ((eps - pred)**2).mean()
# generation: noise in the latent → denoise → decoder(z0) → pixels
pixels latent diffusionin the latent (cheap) latent pixels ED (decoder)
The autoencoder compresses into a latent, the diffusion works there (orders of magnitude cheaper), the decoder brings the pixels back. Text steers it through cross-attention.
Analogy. Why rough out a drawing at full size on a huge canvas when you can sketch it small and scale it up afterwards? The latent is the sketch: all the creative work (the diffusion) happens on a compact version, and the final "printer" (the decoder) blows it up to full resolution. Cheap where expensive is not needed.

Why it matters

This paper is the research foundation of Stable Diffusion (the open-weights release of August 2022): text-to-image started running on CONSUMER GPUs and went mainstream; the openness spawned a huge ecosystem (LoRA for SD, ControlNet and the rest).

Connections

The diffusion math (ε-prediction, the MSE loss) is taken from DDPM unchanged — LDM only moves it into a latent space. It is an engineering optimization on top of DDPM's theoretical foundation.

← controlled via40. CLIP

The text conditioning leans on CLIP-style embeddings — a shared language↔vision space. Cross-attention over those embeddings is the mechanism by which the prompt "steers" the generation.

Cross-attention is Transformer attention applied between the latent and the text. The same "retrieve by relevance" mechanism (#22, #32) now ties the image to the words of the prompt.

Questions worth asking

Does quality not suffer from working in a compressed latent?

Almost not at all — and that is the trick. The autoencoder is trained with a perceptual loss so that it keeps the semantically important detail and throws away only the pixel redundancy nobody perceives. The diffusion works on what matters to perception. Compress too aggressively and it does start to hurt, so the compression factor is an important hyperparameter in the cheap/quality balance.

Why was the open release of Stable Diffusion specifically such an event?

DALL·E 2 (closed) showed that text-to-image was possible; Stable Diffusion made it available — weights open, the model fits on a single consumer GPU. That shifted the power from labs to the community: thousands of fine-tunes, tools (ControlNet), businesses. Openness plus the cheapness of latent diffusion equals a reach the closed equivalents never had.

Is cross-attention conditioning the only way to steer the generation?

No, it is the base layer, but the ecosystem added plenty on top: classifier-free guidance (strengthens the prompt's influence), ControlNet (control over pose and outlines), inpainting masks, img2img. Cross-attention to text is the main channel, but "how do I make the model draw exactly what I want" has grown into an engineering field of its own above it.

What to read in the original

Read the key parts — the latent idea (two-phase training) and the cross-attention conditioning.