Era 5 · The LLM era · 2020

37 Scaling Laws

Scaling Laws for Neural Language Models · Kaplan, McCandlish et al. · OpenAI
🟧 read selectively~1.5 horiginal ↗
The gist in 20 seconds. The loss of an LM falls as a smooth POWER law in model size, data and compute — predictably over 7+ orders of magnitude. The shape of the network barely matters; scale rules. It turned "let's make it bigger" into a predictable engineering discipline and justified GPT-3.

Context

You can scale models up, but how exactly does quality grow, and where should the compute go? Kaplan, McCandlish et al. measure it systematically.

The idea and the mechanism

Having trained many models of different sizes on different amounts of data, they found: the loss falls as a smooth power law in the number of parameters N, the amount of data D and the compute C — predictably over many orders of magnitude. The shape of the network (the width-to-depth ratio) hardly matters.

statistics The power law and the compute-optimal split

The dependence of the test loss on each resource (with the others in abundance) is a power law:

L(N) ≈ (NcN)αN

In log-log coordinates that is a straight line with slope −α — hence the surprising predictability: measure the loss on small models and you can extrapolate to large ones. The practical consequence for splitting a budget: at a fixed C there is an optimal pair (N, D) that minimizes the loss. Kaplan concluded that you are better off taking very large models and not training them to convergence (stopping early). Chinchilla (#42) would later refine that recipe: N and D should be grown more evenly.

NumPy Fitting a power law
import numpy as np
# L ≈ (N_c/N)^α  →  log L is linear in log N
slope, intercept = np.polyfit(np.log(N), np.log(L), 1)
alpha = -slope                  # slope of the log-log line = −α
# extrapolation: the loss for a model 10× larger
L_pred = np.exp(intercept + slope * np.log(10 * N))
log(parameters N) log(loss) a straight line → predictable
In log-log coordinates the loss lands on a straight line: measure the small models and you can extrapolate the behavior of the large ones.
Analogy. The way engineers derive a "law" for the strength of a material from a handful of tests and then design a bridge with confidence, without breaking it first. Scaling laws are strength-of-materials for neural networks: measure a couple of models and you know what quality the giant will reach, before spending millions on training it.

Why it matters

It turned "let's make the model bigger" into predictable engineering and directly justified building GPT-3 a few months later. "The loss is predictable as a function of scale" is the foundation of the whole strategy of the frontier labs: you can plan progress rather than hope for it.

Connections

← formalizes36. GPT-2

GPT-2 showed empirically that "bigger = qualitatively better". Scaling laws turned that observation into a quantitative law with exponents — from an intuition into a planning tool.

→ justifies38. GPT-3

With the law in hand, OpenAI could commit to a 175B model knowing in advance roughly what loss they would get. GPT-3 is a direct consequence of believing that scale is predictable.

↔ refined in42. Chinchilla

Chinchilla re-examined the compute split and found that Kaplan had underestimated the role of data: parameters and tokens should be grown more evenly (~20 tokens per parameter). The same power-law idea, a different conclusion about the optimal point.

Questions worth asking

Why should the loss be a smooth power law at all — is it magic?

Not magic, but not fully understood either. There are theories (via the spectrum of the data, effective dimension, manifold approximation) that account for power-law behavior, but there is no single derivation "from first principles". Empirically the law is startlingly robust over 7+ orders of magnitude — and it is that robustness, not any theoretical inevitability, that made it a working tool.

If the loss is predictable, are abilities (reasoning, code) predictable too?

No, and that is the key caveat. The loss falls smoothly, but particular abilities often appear in a jump at some scale ("emergence", cf. Chain-of-Thought #43). A smooth loss ≠ smooth growth in skills. So scaling laws predict the average loss well, but predict badly when exactly reasoning or arithmetic will "switch on".

Will the law break eventually — is there a ceiling?

Yes: the power law has an "irreducible" term (the entropy of language) — the loss does not fall below the noise floor of the data. On top of that we run into physical limits: high-quality text data is finite, compute is expensive. So the frontier has moved from "simply bigger" towards efficiency (Chinchilla, distillation, better data) and towards new axes of scaling (inference-time compute, as in the reasoning models).

What to read in the original

Read the essentials — the power laws and the conclusion about splitting compute (corrected for Chinchilla).