Era 5 · The LLM era · 2022

42 Chinchilla

Training Compute-Optimal Large Language Models · Hoffmann, Borgeaud et al. · DeepMind · NeurIPS
🟧 read selectively~1.5 horiginal ↗
The gist in 20 seconds. Most LLMs were badly undertrained: for a fixed compute budget, parameters and tokens should grow TOGETHER (~20 tokens per parameter) rather than inflating size alone. Proof: Chinchilla 70B (1.4T tokens) beats Gopher 280B at the same compute.

Context

After the scaling laws (#37) the race was about model SIZE (Gopher 280B, GPT-3 175B, MT-NLG 530B). Hoffmann et al. redo the measurement across 400+ models and find the industry had been scaling the wrong way.

The idea and the mechanism

At a fixed compute budget the number of parameters N and the number of tokens D should grow ROUGHLY IN STEP — the optimum is ~20 tokens per parameter. Most of the giants were undertrained: models too large on data too small.

optimization Compute-optimal: minimizing loss under a fixed budget

Fit a parametric form for the loss (E is the irreducible entropy of language):

L(N, D) = E + ANα + BDβ

The compute budget (FLOPs) is roughly C ≈ 6 N D. Minimize L subject to fixed C (Lagrange multipliers) and the optimum gives N* ∝ Ca, D* ∝ Cb with a ≈ b ≈ 0.5, close to each other:

N and D grow ~in step  ⟹   D/N ≈ 20 tokens/parameter

Kaplan (#37) underestimated the role of D and recommended "large models, few tokens". The reason for the discrepancy is concrete: Kaplan's learning-rate schedule was not run to completion for the token budget (which understated the gain from larger D), and embedding parameters were counted differently — hence the skewed exponents. Chinchilla, fixing both the loss form and the LR schedule, showed that data must grow alongside parameters. The proof is Chinchilla (70B, 1.4T) beating Gopher (280B, 300B tokens) at the same C.

Python The Chinchilla rule
# budget C ≈ 6·N·D FLOPs; the optimum is to grow N and D in step
def chinchilla_tokens(N):
    return 20 * N            # ~20 tokens per parameter

# Gopher: 280B parameters, 300B tokens  →  D/N ≈ 1  (badly undertrained!)
# Chinchilla: 70B parameters, 1.4T tokens →  D/N ≈ 20  ✓
Gopher280Bfew tokens Chinchilla70B, 1.4T Chinchilla < Gopher (smaller, but trainedlonger → lower loss) same compute
At equal compute the smaller but longer-trained Chinchilla beats a Gopher four times its size — the data mattered more than the size.
Analogy. Hiring one genius and never letting them learn anything properly is worse than taking a capable person and training them well. Inflating the model (hiring geniuses) while skimping on data (on the training) is money down the drain. Chinchilla showed that on a fixed budget you are better off with "medium size + plenty of schooling".

Why it matters

It turned the industry towards data-heavy training (hence the trillions of tokens in LLaMA and everything after) and redefined "compute-optimal"; the ~20:1 rule (the "Chinchilla point") became a design heuristic.

Connections

↔ refines37. Scaling Laws

The same power-law idea, a different conclusion about where the optimum sits. Kaplan said "large models, few tokens"; Chinchilla says "grow N and D in step". A correction that overturned frontier practice.

→ steers50. LLaMA

LLaMA is a direct application of the Chinchilla lesson: take smaller models (7–65B) but train them on far more tokens, for the sake of cheap inference. A 13B LLaMA beats GPT-3 175B precisely because it was trained longer.

↔ context38. GPT-3

GPT-3 (175B on ~300B tokens) is the archetypal "undertrained giant" by Chinchilla's standards. With hindsight: the same model or better could have been had at a smaller size, trained on more tokens.

Questions worth asking

If the Chinchilla optimum is so good, why are models still built large?

Because "compute-optimal for training" ≠ "optimal for deployment". A small model is cheaper at inference, so it is often worth training it past the Chinchilla point (more tokens than 20×N): you spend more on training and save it back over millions of requests. That is exactly what LLaMA does. Chinchilla optimizes training; the real economics looks at inference.

Is 20 tokens per parameter a universal constant?

No — it is an estimate for one particular loss form, architecture and schedule of that period. The exact ratio depends on the data, the tokenizer and the training details; later work reports slightly different numbers. "~20" is a useful heuristic, not a law of nature. The main conclusion — grow data alongside parameters — is far more robust than the constant itself.

With that appetite for tokens, will we run out of data?

Yes, that is a real ceiling. There is a finite amount of good text on the internet, and the frontier is already approaching the end of it. Hence the interest in synthetic data, multiple passes over the corpus, multimodality and new axes of scaling (inference-time compute). "Data matters more than we thought" unexpectedly turns it into a scarce resource.

What to read in the original

Read the essentials — the equal-scaling law and the Chinchilla vs Gopher result.