Era 6 · Generative models and systems · 2024

58 o1 / Test-Time Compute

Learning to Reason with LLMs (OpenAI o1, 2024) · Scaling LLM Test-Time Compute Optimally… · Snell et al., 2024
🟥 read in full~2 horiginal ↗
The gist in 20 seconds. A new axis of scaling. Until now what grew was training (Kaplan/Chinchilla). o1 (Sep 2024) showed that RL-trained long reasoning plus more "thinking" at inference sharply lifts quality on hard problems — test-time compute became a scaling axis. Snell et al.: with a compute-optimal allocation of the inference budget you can be ×4 more efficient than best-of-N, and a small model plus test-time compute can beat a big one.

Context

Scaling laws (#37) and Chinchilla (#42) scaled training: more parameters and more tokens. But good data is finite (#42), and leaning on a single lever eventually runs out. Another lever was left almost untouched — spending compute not on training but on inference: letting the model "think" longer about its answer.

The idea and the mechanism

o1 is trained with RL to produce a long internal chain of reasoning — not as a prompting trick (CoT #43) but as a learned ability: writing out steps, checking itself, going back over mistakes, trying different routes. The key observation: quality grows monotonically both with train-time RL and with test-time "thinking" — the more reasoning tokens at inference, the more accurate the answer on hard problems. Snell et al. formalize how to spend the test-time budget optimally (search against verifier models vs adaptive revision of the answer) and show that allocating the budget by prompt difficulty (compute-optimal) is up to ×4 more efficient than naive best-of-N — and that sometimes a small model with test-time compute beats a far bigger one.

optimization · scaling Two budget axes and verifier search

The simplest way to spend test-time compute is to sample N answers and pick the best one by a verifier model V:

\[ y^{\star} = \arg\max_{y \in \{y_1,\dots,y_N\}} V(y),\qquad y_i \sim \pi(\cdot \mid x) \]

There is also a "vertical" way — making one chain longer and better (adaptive revision). Quality is a function of two budgets; for a fixed total, the optimal split shifts towards inference on hard problems:

\[ \text{quality} = f\big(C_{\text{train}},\, C_{\text{test}}\big),\qquad \text{fix } C_{\text{total}} \;\Rightarrow\; \text{shift budget to } C_{\text{test}} \text{ on hard } x \]

Snell et al.: naive best-of-N spends the same budget on easy and hard prompts alike; a compute-optimal strategy allocates it adaptively by difficulty — hence the >4× efficiency win. RL (o1), meanwhile, teaches the model to spend its "thinking" usefully rather than merely longer.

Python Test-time scaling: best-of-N with a verifier
def answer(model, verifier, x, N):
    ys = [model.sample(x) for _ in range(N)]      # N chains of reasoning — this is the test-time compute we spend
    return max(ys, key=verifier.score)             # pick the best one by the verifier
# compute-optimal: N depends on the difficulty of x (small for easy, large for hard)
# o1 goes further: RL teaches the model to produce ONE long self-checking chain
test-time compute ("thinking") → quality → o1: thinks longer → more accurate ordinary model (no reasoning RL) the gain is largest on hard problems
Accuracy grows with the time spent "thinking" at inference — a new scaling axis alongside training. RL teaches the model to spend that thinking usefully (self-checking, backtracking) rather than simply generating for longer.
Analogy. An exam. You can hire a more capable person (grow the training — parameters and data). Or you can give the same person more time to think, write the solution out and check it over. On hard problems the second is often cheaper and more effective than hunting for a genius. And o1 is also a trained ability to use that time: not just sitting there longer, but thinking in a structured way.

Why it matters

It shifted the frontier's paradigm from "more parameters" to "think more" and opened up a class of reasoning models (o1, then o3, DeepSeek-R1 and a wave of others). It is a direct answer to running out of data (#42): as the training axis gets harder to scale, a second one appears — inference. For system design it changes the economics: part of the "intelligence" moves out of the model weights and into the inference budget.

Connections

← a new axis on42. Chinchilla

Chinchilla optimized the training compute (parameters vs tokens). Test-time compute adds an orthogonal axis: spending at inference. When training data is scarce, it is the inference axis that gives the next jump.

CoT (2022) showed that reasoning step by step helps — but through the prompt. o1 turns the length and quality of that reasoning into a resource you can scale accuracy with, and teaches it by RL rather than by prompting.

↔ open implementation53. DeepSeek V3 / R1

R1 is the public counterpart of the o1 idea: reasoning grown by RL on verifiable rewards. The same line of "test-time reasoning as an ability", but with open weights and a written-down recipe (see #59 GRPO).

Questions worth asking

How is o1's "thinking longer" different from just a long CoT prompt?

CoT (#43) is a prompting trick: we ask the model to reason step by step. o1 is trained (by RL) to generate long chains that are actually useful — to check itself, drop dead-end branches, come back to mistakes. Thinking becomes a learnable ability, and its length becomes a quality dial. The difference is like the one between asking someone to think out loud and training them to think effectively.

Why can a small model plus test-time compute beat a big one?

On hard problems, search, verification and revision at inference add "effective capacity" more cheaply than inflating the parameter count. Snell et al. showed that allocating the test-time budget by difficulty (compute-optimal) is >4× more efficient than best-of-N — and a small model given time to "think" beats a big one that answers straight away. Compute moves from training into inference.

Is there a limit to test-time scaling?

Yes. On easy problems the extra thinking does not help (and sometimes "overthinking" / overcomplicating actively hurts). Inference costs rise — answers become expensive and slow. The gain depends on the quality of the verifier or the reward. And it does not replace training, it complements it. Test-time compute is a powerful lever, but not a free or a universal one.

What to read in the original

Read both in full. OpenAI's "Learning to Reason with LLMs" — the curves of quality growing with both train-time RL and test-time thinking (the central claim). Snell et al. (arXiv 2408.03314) — the formalization of compute-optimal test-time: search against a verifier vs revision, allocating the budget by difficulty, and the comparison with best-of-N.