Era 5 · The LLM era · 2018

35 GPT-1

Improving Language Understanding by Generative Pre-Training · Radford, Narasimhan, Salimans, Sutskever · OpenAI
🟦 this write-up is enough~30 minoriginal ↗
The gist in 20 seconds. A Transformer decoder, generatively pre-trained on unlabelled text and then fine-tuned per task, beats purpose-built architectures. The start of the GPT line: decoder-only, autoregressive, generative — as opposed to the bidirectional BERT.

Context

The same 2018; OpenAI tries a recipe running parallel to BERT — generative pre-training of a Transformer decoder on a stream of text.

The idea and the mechanism

Stage 1 (unsupervised): train an ordinary language model (next-token prediction) on a large corpus of books. Stage 2 (supervised): for each task, reformat the input slightly, add a linear head and fine-tune on labeled data. Minimal task-specific architecture — nearly all the knowledge sits in the pre-trained decoder.

probability Two stages: pre-training and fine-tuning

Stage 1 — generative pre-training. Maximize the likelihood of the next token (a causal LM) on unlabelled text:

L1 = Σt log P(xt | xt−k..t−1; θ)

Stage 2 — supervised fine-tuning. For labeled (x, y) add a linear head on top of the final state and maximize:

L2 = Σ log P(y | x1..m)  (+ λ L1 as an auxiliary term)

The key: the heavy lifting — knowledge of the language — happens in stage 1, out of cheap unlabelled text, and the expensive labeled data is only needed to "tune" the model to the task in stage 2. This is transfer learning for NLP, which will later be pushed all the way to zero-shot (GPT-2/3).

PyTorch The causal LM loss (pre-training)
import torch.nn.functional as F
# predict the next token: shift the sequence by 1
def lm_loss(logits, tokens, V):
    return F.cross_entropy(logits[:, :-1].reshape(-1, V),
                           tokens[:, 1:].reshape(-1))   # target = input shifted left
pre-trainingunlabelled text fine-tuninglabeled task answer learn the language (cheap, plentiful) tune to the task (expensive, scarce)
Cheap pre-training on text supplies the "knowledge of the language"; expensive labels are only needed for the final tuning to a particular task.
Analogy. First a person simply reads a great deal — building up general literacy and background knowledge (pre-training). Then, taking a specific job, they do a short specialist course (fine-tuning). Learning the language again from scratch for every job would be daft — you build the general base once and reuse it.

Why it matters

It laid down the GPT line (decoder-only, autoregressive, unidirectional) — as opposed to the bidirectional encoder of BERT. That line leads to GPT-2/3 and ChatGPT. It came out quietly, as a PDF report with no conference; conceptually it was superseded by GPT-2/3.

Connections

← builds on32. Transformer

GPT is the decoder half of the Transformer put through large-scale pre-training. Causal self-attention (it sees only the past) is exactly what autoregressive generation needs, unlike BERT's bidirectional encoder.

→ scales into36. GPT-2

GPT-2 takes the same decoder-only scheme and simply makes the model and the data much bigger — and discovers zero-shot abilities. A direct continuation: same recipe, different scale.

↔ contrast34. BERT

The twins of 2018, with opposite philosophies: BERT does bidirectional understanding (fine-tuned per task), GPT does autoregressive generation (later steered by the prompt). Their fork defined two classes of LLM.

Questions worth asking

Why did the GPT line eventually "win" over BERT in the public mind?

Because generation scales into a general-purpose assistant driven by a prompt, with no fine-tuning per task — and that is an interface anyone understands (ChatGPT). BERT remained a powerful understanding component inside systems (search, classification), but it does not "talk". What won was not the architecture as such but the paradigm of "one generative model for everything, through text".

Why train on books (BooksCorpus) specifically, rather than any text?

Books give you long coherent passages — the model learns long-range dependencies rather than fragments. That matters for generating coherent text. Later (GPT-2/3) the corpus was widened to quality web text, but the idea that "long coherent context is better for an LM" stuck.

Fine-tuning per task — is that not exactly what the field later moved away from?

Precisely. GPT-1 still needed supervised fine-tuning in stage 2. GPT-2 showed that at sufficient scale you can solve tasks zero-shot through the prompt, and GPT-3 few-shot. In two steps the GPT line moved from "fine-tune it for the task" to "describe the task in the prompt" — and that is the main shift of paradigm.

What to read in the original

This write-up is enough — the concept was superseded by GPT-2/3. If you are curious, look at how minimally the input is reformatted for the different tasks in stage 2.