35 GPT-1
Context
The same 2018; OpenAI tries a recipe running parallel to BERT — generative pre-training of a Transformer decoder on a stream of text.
The idea and the mechanism
Stage 1 (unsupervised): train an ordinary language model (next-token prediction) on a large corpus of books. Stage 2 (supervised): for each task, reformat the input slightly, add a linear head and fine-tune on labeled data. Minimal task-specific architecture — nearly all the knowledge sits in the pre-trained decoder.
probability Two stages: pre-training and fine-tuning
Stage 1 — generative pre-training. Maximize the likelihood of the next token (a causal LM) on unlabelled text:
Stage 2 — supervised fine-tuning. For labeled (x, y) add a linear head on top of the final state and maximize:
The key: the heavy lifting — knowledge of the language — happens in stage 1, out of cheap unlabelled text, and the expensive labeled data is only needed to "tune" the model to the task in stage 2. This is transfer learning for NLP, which will later be pushed all the way to zero-shot (GPT-2/3).
PyTorch The causal LM loss (pre-training)
import torch.nn.functional as F
# predict the next token: shift the sequence by 1
def lm_loss(logits, tokens, V):
return F.cross_entropy(logits[:, :-1].reshape(-1, V),
tokens[:, 1:].reshape(-1)) # target = input shifted left
Why it matters
It laid down the GPT line (decoder-only, autoregressive, unidirectional) — as opposed to the bidirectional encoder of BERT. That line leads to GPT-2/3 and ChatGPT. It came out quietly, as a PDF report with no conference; conceptually it was superseded by GPT-2/3.
Connections
GPT is the decoder half of the Transformer put through large-scale pre-training. Causal self-attention (it sees only the past) is exactly what autoregressive generation needs, unlike BERT's bidirectional encoder.
GPT-2 takes the same decoder-only scheme and simply makes the model and the data much bigger — and discovers zero-shot abilities. A direct continuation: same recipe, different scale.
The twins of 2018, with opposite philosophies: BERT does bidirectional understanding (fine-tuned per task), GPT does autoregressive generation (later steered by the prompt). Their fork defined two classes of LLM.
Questions worth asking
Why did the GPT line eventually "win" over BERT in the public mind?
Because generation scales into a general-purpose assistant driven by a prompt, with no fine-tuning per task — and that is an interface anyone understands (ChatGPT). BERT remained a powerful understanding component inside systems (search, classification), but it does not "talk". What won was not the architecture as such but the paradigm of "one generative model for everything, through text".
Why train on books (BooksCorpus) specifically, rather than any text?
Books give you long coherent passages — the model learns long-range dependencies rather than fragments. That matters for generating coherent text. Later (GPT-2/3) the corpus was widened to quality web text, but the idea that "long coherent context is better for an LM" stuck.
Fine-tuning per task — is that not exactly what the field later moved away from?
Precisely. GPT-1 still needed supervised fine-tuning in stage 2. GPT-2 showed that at sufficient scale you can solve tasks zero-shot through the prompt, and GPT-3 few-shot. In two steps the GPT line moved from "fine-tune it for the task" to "describe the task in the prompt" — and that is the main shift of paradigm.
What to read in the original
This write-up is enough — the concept was superseded by GPT-2/3. If you are curious, look at how minimally the input is reformatted for the different tasks in stage 2.