34 BERT
Context
Before 2018 pre-training in NLP was either unidirectional (a left-to-right LM, as in GPT) or shallow (word2vec). Devlin et al. make it deep and BIdirectional.
The idea and the mechanism
A Transformer ENCODER (it sees the whole sequence at once). Pre-training on two self-supervised tasks: Masked LM — mask ~15% of the tokens and predict them from the context on BOTH sides; Next Sentence Prediction — do these two sentences follow each other (later judged to be of little use). Then, for a given task, you bolt on a small "head" and fine-tune the whole model.
probability Masked LM vs an ordinary language model
An ordinary (causal) LM models P(xt | x<t) — the left context only; a token's representation knows nothing about what is to its right. A masked LM masks a subset of positions M and predicts them from everything else:
Since x\M includes tokens both to the left and to the right of i, every token's representation becomes genuinely bidirectional — something a left-to-right LM cannot achieve (the token would "see" its own answer). The price: the [MASK] token exists only during training and is absent at inference (a train/test mismatch). BERT's signature trick is to split the chosen 15%: 80% → [MASK], 10% → a random token, 10% → leave the original. The substitutions and the untouched originals force the model to build a good representation for every position (it never knows whether the token in front of it is genuine), which softens the gap with inference.
PyTorch The masked LM loss
import torch.nn.functional as F
# labels = −100 at unmasked positions → ignored in the loss
def mlm_loss(logits, labels, V):
return F.cross_entropy(logits.view(-1, V), labels.view(-1),
ignore_index=-100) # masked positions only
Why it matters
A jump on 11 NLP benchmarks at once (GLUE +7.7), and "pre-train → fine-tune" became the standard. The encoder style (understanding: classification, NER, retrieval, extractive QA) is the branch running parallel to the generative GPTs; embeddings from the BERT family are still at work in search and RAG.
Connections
BERT is the encoder half of the Transformer, put through large-scale self-supervised pre-training. Without self-attention, bidirectionality would have been inefficient; the Transformer supplied both the architecture and the parallelism to train on enormous corpora.
Two branches of one idea: "pre-train a Transformer on text". BERT is encoder-only, bidirectional, built for understanding; GPT is decoder-only, autoregressive, built for generation. They parted ways in 2018 and defined two classes of model.
The idea of "pre-train an encoder on a self-supervised signal, then use the representations" generalizes beyond text. CLIP does the same for image-text pairs, and its text encoder is a direct descendant of the BERT approach.
Questions worth asking
Why is BERT almost never used to generate text?
It is bidirectional and trained to fill in gaps, not to continue text left to right. Autoregressive generation needs a causal model (GPT) that at each step sees only the past. BERT is brilliant at understanding (classification, search, extraction), but by construction it is not built to produce coherent text.
Only 15% of tokens get masked — why not more, why not less?
A trade-off. Mask too little and you get little signal per pass, so training is inefficient. Mask too much and too little context remains: the task becomes unsolvable and the representations suffer. 15% empirically balances "enough targets" against "enough context". Later work (masking 40% in large models, for instance) showed the optimum depends on model size.
Why was NSP (next sentence prediction) later dropped?
Follow-up work (RoBERTa) showed NSP barely helps and sometimes hurts — the task is too easy (telling contiguous from random sentences is trivial from topic alone). Dropping NSP and training MLM longer on a bigger corpus, RoBERTa beat BERT. A good lesson: not every "sensible" auxiliary task is actually useful — you have to check.
What to read in the original
Read the essentials — the MLM and the fine-tuning scheme; the ablations and the NSP details can be skimmed (bearing in mind NSP was later judged redundant).