38 GPT-3
Context
GPT-2 hinted that scale buys you zero-shot; the scaling laws said how to scale. OpenAI builds a 175B-parameter model.
The idea and the mechanism
The main discovery is in-context few-shot learning: give the model a few examples of a task in the PROMPT (input→output) and it performs a new task simply by picking up the pattern from the context, WITHOUT any weight update. One frozen model solves a mass of tasks through a text prompt alone (zero/one/few-shot).
probability In-context learning without a weight update
Few-shot is just a conditional probability from a frozen model, with the examples slipped into the condition:
The key difference from classic training: the weights do not change. The "learning" happens entirely in the forward pass — attention reads the regularity off the examples in the context and applies it to the query. There is no gradient step at all. This is a qualitatively new phenomenon: one model adapts to a task from a handful of examples on the fly, emerging with scale (much more pronounced at 175B than at smaller sizes).
Python Few-shot through the prompt (no fine-tuning)
# k examples straight in the prompt; we do not touch the model weights
prompt = """2 + 2 = 4
7 + 5 = 12
3 + 9 ="""
out = lm.generate(prompt) # the model continues: " 12"
# same model, different prompt → different task
Why it matters
It shifted the whole usage paradigm — from "train a model for the task" to "describe the task in the prompt" (in-context learning, prompt engineering). It laid down the product model GPT-3 API → ChatGPT. The weaknesses (honestly flagged in the paper): arithmetic, long-range consistency, factuality, sensitivity to how the prompt is worded.
Connections
GPT-2 showed zero-shot at 1.5B; GPT-3 pushed the scale to 175B and uncovered few-shot in-context learning — a qualitatively new capability that emerged with scale. The same decoder-only design, two more orders of magnitude.
GPT-3 is powerful but raw — it continues text rather than following instructions. RLHF aligns it with human intent and turns it into an assistant (ChatGPT). The few-shot model is the foundation; alignment is what made it useful to the wider world.
Few-shot prompting of GPT-3 is the interface through which Chain-of-Thought would later be discovered: put reasoning steps into the examples and you sharply raise the model's arithmetic and logic. CoT is "advanced few-shot", grown straight out of this paradigm.
Questions worth asking
The model "learns" from examples without changing weights — is that learning at all?
Strictly speaking, no: the parameters stay fixed. It is conditional inference — attention finds the regularity in the examples and extrapolates it. There are hypotheses that inside the forward pass the model implicitly performs something like an optimization step ("mesa-optimization"), but that is contested. It is more accurate to call in-context learning adaptation from context rather than learning in the usual sense.
Few-shot is sensitive to wording and to the order of examples — how reliable is it?
Brittle. Reordering the examples, their format, even the choice of separators can noticeably change the answer; sometimes "random" labels in the examples work almost as well as correct ones (the model latches onto the format, not only the content). This spawned a whole discipline of prompt engineering. Few-shot is a powerful but temperamental interface, and its reliability is an engineering concern in its own right.
Were 175B parameters "too many", or a natural next step?
By the scaling laws, natural: the loss fell predictably and 175B delivered the expected gain. But few-shot as a qualitative capability came as a surprise — the smooth loss curve did not predict it. Chinchilla would later show that 175B was undertrained for that amount of data: a smaller model on more tokens would have been the better buy. So: "natural by the loss, suboptimal by the recipe".
What to read in the original
Read the essentials — the intro, the few-shot mechanism, the limitations and broader impacts sections; the 40+ pages of per-task tables are reference material.