36 GPT-2
Context
GPT-1 was fine-tuned per task. OpenAI checks: what if you simply make the model and the data MUCH bigger?
The idea and the mechanism
The same decoder-only Transformer, but up to 1.5B parameters, trained on WebText (a curated web corpus). No mechanism changes — scale only. And it turns out: the model does translation, question answering and summarization ZERO-SHOT, with no fine-tuning, as long as the task is expressed in the text of the prompt.
probability Why an ordinary LM solves tasks zero-shot
A language model learns one thing — the distribution P(x) over text. But if the task and the input are expressed as text, the conditional probability of the continuation is the answer:
Given "Translate to French: cat ⟹", for instance, the model continues with "chat". Where does that skill come from? A web corpus contains masses of implicit demonstrations of tasks — translations, questions and answers, summaries, code with comments. In minimizing the language-modeling loss on such text, the model incidentally learns to perform those tasks. Hence the title: "unsupervised multitask learners" — multitasking arises by itself, as a by-product of next-word prediction on varied data.
Python Zero-shot through the prompt
# the task is stated in the prompt itself, no fine-tuning and no examples
prompt = "Translate to French: cat =>"
out = lm.generate(prompt) # the model continues: " chat"
prompt = "Summary: <long text> TL;DR:"
out = lm.generate(prompt) # the model continues with a summary
Why it matters
A shift towards the idea of a general-purpose LM steered by a prompt — which GPT-3 would confirm outright, and which would become ChatGPT. Famous for its staged release over fears of misuse ("too dangerous to release" → later "no evidence of misuse").
Connections
Same architecture, same recipe — only more model and more data. GPT-2 is the experiment "what does pure scale buy you", and the answer turned out to be a qualitative jump (zero-shot), not merely a quantitative one.
GPT-2 hinted at zero-shot; GPT-3 pushed the scale two orders of magnitude further and opened up few-shot in-context learning. A straight ladder of scale: 1.5B → 175B, and with it the move from "a hint of abilities" to "a general few-shot solver".
GPT-2's observation that "bigger = qualitatively better" was soon formalized as scaling laws: the loss falls as a predictable power law in scale. That turned "let's make it bigger" from an intuition into an engineering strategy.
Questions worth asking
"Too dangerous to release" — a real threat, or marketing?
Contested. OpenAI justified the staged release by the risk of mass-produced disinformation; critics called it inflated and a PR move. In the event no catastrophe happened, and the full model was published about nine months later. The episode matters as the first loud precedent in the debate about "responsible release" of powerful models — a topic that has only sharpened since.
Zero-shot works, but badly — why is it still counted as a breakthrough?
Because what matters is not the quality level but the fact: a model that nobody taught to translate translates — purely out of language modeling. That changed the question from "how do we train it for the task" to "how does scale give rise to abilities". The weak zero-shot of 2019 is the forerunner of GPT-3's strong few-shot and of instruct models; the breakthrough is in the direction, not in the metric.
Corpus quality (WebText) — how critical is it?
Very. WebText was assembled from links posted on Reddit with positive karma — a crude but effective quality filter. Junk web text would have made the model worse; curating the data turned out to matter as much as scale. The lesson that "data decides" has only grown stronger since — modern frontier models spend enormous effort on filtering and corpus composition.
What to read in the original
Read the essentials — the zero-shot framing and the role of WebText; the benchmark tables can be skimmed.