Era 5 · The LLM era · 2020

38 GPT-3

Language Models are Few-Shot Learners · Brown, Mann, Ryder et al. (31 authors) · OpenAI · NeurIPS
🟧 read selectively~2 horiginal ↗
The gist in 20 seconds. A 175B-parameter LM with a breakthrough capability — in-context few-shot learning: put a couple of examples in the prompt and the model performs a new task WITHOUT any weight update. One frozen model handles a mass of tasks through text alone. It shifted the paradigm towards prompting.

Context

GPT-2 hinted that scale buys you zero-shot; the scaling laws said how to scale. OpenAI builds a 175B-parameter model.

The idea and the mechanism

The main discovery is in-context few-shot learning: give the model a few examples of a task in the PROMPT (input→output) and it performs a new task simply by picking up the pattern from the context, WITHOUT any weight update. One frozen model solves a mass of tasks through a text prompt alone (zero/one/few-shot).

probability In-context learning without a weight update

Few-shot is just a conditional probability from a frozen model, with the examples slipped into the condition:

P(y | (x1,y1), …, (xk,yk), xquery)

The key difference from classic training: the weights do not change. The "learning" happens entirely in the forward pass — attention reads the regularity off the examples in the context and applies it to the query. There is no gradient step at all. This is a qualitatively new phenomenon: one model adapts to a task from a handful of examples on the fly, emerging with scale (much more pronounced at 175B than at smaller sizes).

Python Few-shot through the prompt (no fine-tuning)
# k examples straight in the prompt; we do not touch the model weights
prompt = """2 + 2 = 4
7 + 5 = 12
3 + 9 ="""
out = lm.generate(prompt)     # the model continues: " 12"
# same model, different prompt → different task
examples: cat ⟹ chatdog ⟹ chien query: house ⟹ ? LM (frozen) maison no training — context only
Few-shot: the examples sit in the prompt and the model produces the answer to a new query. The weights do not change — the "learning" happens in the forward pass.
Analogy. A quick-witted person you show two or three rounds of an unfamiliar game and then tell "your turn". They do not "retrain their brain" — they grasp the rule from the examples and apply it immediately. GPT-3 does the same: examples in the prompt are "show, don't explain", and the model picks the pattern up on the fly.

Why it matters

It shifted the whole usage paradigm — from "train a model for the task" to "describe the task in the prompt" (in-context learning, prompt engineering). It laid down the product model GPT-3 API → ChatGPT. The weaknesses (honestly flagged in the paper): arithmetic, long-range consistency, factuality, sensitivity to how the prompt is worded.

Connections

← scales up36. GPT-2

GPT-2 showed zero-shot at 1.5B; GPT-3 pushed the scale to 175B and uncovered few-shot in-context learning — a qualitatively new capability that emerged with scale. The same decoder-only design, two more orders of magnitude.

GPT-3 is powerful but raw — it continues text rather than following instructions. RLHF aligns it with human intent and turns it into an assistant (ChatGPT). The few-shot model is the foundation; alignment is what made it useful to the wider world.

→ unlocked by43. Chain-of-Thought

Few-shot prompting of GPT-3 is the interface through which Chain-of-Thought would later be discovered: put reasoning steps into the examples and you sharply raise the model's arithmetic and logic. CoT is "advanced few-shot", grown straight out of this paradigm.

Questions worth asking

The model "learns" from examples without changing weights — is that learning at all?

Strictly speaking, no: the parameters stay fixed. It is conditional inference — attention finds the regularity in the examples and extrapolates it. There are hypotheses that inside the forward pass the model implicitly performs something like an optimization step ("mesa-optimization"), but that is contested. It is more accurate to call in-context learning adaptation from context rather than learning in the usual sense.

Few-shot is sensitive to wording and to the order of examples — how reliable is it?

Brittle. Reordering the examples, their format, even the choice of separators can noticeably change the answer; sometimes "random" labels in the examples work almost as well as correct ones (the model latches onto the format, not only the content). This spawned a whole discipline of prompt engineering. Few-shot is a powerful but temperamental interface, and its reliability is an engineering concern in its own right.

Were 175B parameters "too many", or a natural next step?

By the scaling laws, natural: the loss fell predictably and 175B delivered the expected gain. But few-shot as a qualitative capability came as a surprise — the smooth loss curve did not predict it. Chinchilla would later show that 175B was undertrained for that amount of data: a smaller model on more tokens would have been the better buy. So: "natural by the loss, suboptimal by the recipe".

What to read in the original

Read the essentials — the intro, the few-shot mechanism, the limitations and broader impacts sections; the 40+ pages of per-task tables are reference material.