17 Word2Vec
Context
After Bengio's neural language model (#13) it was clear that embeddings are useful, but training them through a full LM is expensive. Mikolov and his team (Google) make them cheap and at scale.
The idea and the mechanism
Drop the expensive hidden layer and reduce everything to a simple task. CBOW predicts the center word from a window of context; Skip-gram does the reverse, predicting the context words from the word. Training is sped up by negative sampling: instead of normalizing a softmax over the whole vocabulary, tell the real context word apart from a handful of random ones. You can push billions of words through it even on a CPU.
probability Negative sampling: from softmax to binary classification
An honest probability for a context word needs a softmax over the whole vocabulary — a sum over V words at every step, far too expensive:
Negative sampling replaces that with a "real pair or random pair" task: for a genuine pair (w, c) and a few random "negatives" nk we maximize
No sum over the vocabulary — only over a few negatives. Why the analogies appear. The objective pulls together words with similar contexts, and systematic relations ("male→female", "country→capital") end up encoded as the same offset in the space. Hence vking − vman + vwoman ≈ vqueen — the analogy is solved by vector arithmetic.
NumPy An analogy through embedding arithmetic
import numpy as np
# a − b + c ≈ ? (king − man + woman ≈ queen)
def analogy(a, b, c, E, vocab):
v = E[vocab[a]] - E[vocab[b]] + E[vocab[c]]
sims = E @ v / (np.linalg.norm(E, axis=1) * np.linalg.norm(v))
return vocab.itos[sims.argmax()] # nearest by cosine
Why it matters
Dense pretrained embeddings became the standard input to NLP for years. The principle "learn representations through a self-supervised predict-from-context task" is a direct ancestor of BERT/GPT: the same embeddings, only now a Transformer builds the context.
Connections
Bengio trained embeddings as a detail of a language model. Word2vec made them the goal in themselves and threw out the expensive hidden layer for speed — the shift from "embedding as a by-product" → "embedding as the product".
Two routes to the same result. Word2vec is predictive (local windows, negative sampling); GloVe is count-based (a global co-occurrence matrix). Quality is comparable, and GloVe explains why analogy arithmetic works.
The idea "a word is a vector learned from context" is the foundation of the input layer of any LLM. The Transformer adds contextual embeddings (a word's vector depends on the sentence), but static word2vec vectors are their direct ancestor.
Questions worth asking
Are the king−man+woman≈queen analogies a real property, or cherry-picked pretty examples?
A bit of both. The linear structure is real and measurable across many relations, but the analogy metric is sensitive to the details: the words a, b, c are usually excluded from the candidate set, otherwise the answer often turns out to be one of them. Some relations are captured badly. So the effect is genuine, but inflated by a selection of striking examples — healthy scepticism is in order.
What is the fundamental limitation of these embeddings?
They are static: a word has one vector regardless of context. "Bank" (the riverside) and "bank" (the institution) get the same embedding — polysemy is not resolved. That is exactly what contextual models fix (ELMo, BERT, GPT): there a word's vector depends on the whole sentence.
Why does negative sampling work at all — it isn't a real probability?
It is noise-contrastive estimation: instead of modeling the full distribution, you learn to tell data from noise. Provably, with the right choice of noise distribution such a task yields the same embeddings as a full softmax, but without the sum over the vocabulary. You sacrifice exact normalization for speed — and lose almost nothing in representation quality.
What to read in the original
Read the core part — CBOW/Skip-gram. The training details (negative sampling) are covered more fully in the companion paper 1310.4546; the crisp formulation of analogies is in the NAACL paper "Linguistic Regularities".