37 Scaling Laws
Context
You can scale models up, but how exactly does quality grow, and where should the compute go? Kaplan, McCandlish et al. measure it systematically.
The idea and the mechanism
Having trained many models of different sizes on different amounts of data, they found: the loss falls as a smooth power law in the number of parameters N, the amount of data D and the compute C — predictably over many orders of magnitude. The shape of the network (the width-to-depth ratio) hardly matters.
statistics The power law and the compute-optimal split
The dependence of the test loss on each resource (with the others in abundance) is a power law:
In log-log coordinates that is a straight line with slope −α — hence the surprising predictability: measure the loss on small models and you can extrapolate to large ones. The practical consequence for splitting a budget: at a fixed C there is an optimal pair (N, D) that minimizes the loss. Kaplan concluded that you are better off taking very large models and not training them to convergence (stopping early). Chinchilla (#42) would later refine that recipe: N and D should be grown more evenly.
NumPy Fitting a power law
import numpy as np
# L ≈ (N_c/N)^α → log L is linear in log N
slope, intercept = np.polyfit(np.log(N), np.log(L), 1)
alpha = -slope # slope of the log-log line = −α
# extrapolation: the loss for a model 10× larger
L_pred = np.exp(intercept + slope * np.log(10 * N))
Why it matters
It turned "let's make the model bigger" into predictable engineering and directly justified building GPT-3 a few months later. "The loss is predictable as a function of scale" is the foundation of the whole strategy of the frontier labs: you can plan progress rather than hope for it.
Connections
GPT-2 showed empirically that "bigger = qualitatively better". Scaling laws turned that observation into a quantitative law with exponents — from an intuition into a planning tool.
With the law in hand, OpenAI could commit to a 175B model knowing in advance roughly what loss they would get. GPT-3 is a direct consequence of believing that scale is predictable.
Chinchilla re-examined the compute split and found that Kaplan had underestimated the role of data: parameters and tokens should be grown more evenly (~20 tokens per parameter). The same power-law idea, a different conclusion about the optimal point.
Questions worth asking
Why should the loss be a smooth power law at all — is it magic?
Not magic, but not fully understood either. There are theories (via the spectrum of the data, effective dimension, manifold approximation) that account for power-law behavior, but there is no single derivation "from first principles". Empirically the law is startlingly robust over 7+ orders of magnitude — and it is that robustness, not any theoretical inevitability, that made it a working tool.
If the loss is predictable, are abilities (reasoning, code) predictable too?
No, and that is the key caveat. The loss falls smoothly, but particular abilities often appear in a jump at some scale ("emergence", cf. Chain-of-Thought #43). A smooth loss ≠ smooth growth in skills. So scaling laws predict the average loss well, but predict badly when exactly reasoning or arithmetic will "switch on".
Will the law break eventually — is there a ceiling?
Yes: the power law has an "irreducible" term (the entropy of language) — the loss does not fall below the noise floor of the data. On top of that we run into physical limits: high-quality text data is finite, compute is expensive. So the frontier has moved from "simply bigger" towards efficiency (Chinchilla, distillation, better data) and towards new axes of scaling (inference-time compute, as in the reasoning models).
What to read in the original
Read the essentials — the power laws and the conclusion about splitting compute (corrected for Chinchilla).