24 GloVe
Context
By 2014 there are two camps in embeddings: count-based methods (working off the co-occurrence matrix, like LSA) and predictive ones (word2vec). Pennington, Socher and Manning at Stanford put them together.
The idea and the mechanism
We build a global co-occurrence matrix Xij (how often word j appears in the context of i across the whole corpus) and train vectors so that their dot product approximates the logarithm of the count. Quality is on a par with word2vec, but training runs on global statistics rather than local windows.
linear algebra Why ratios of probabilities — and the log-bilinear form
Compare two words through a "probe" word k. A probability on its own, P(k|i), tells you little; their ratio, however, does distinguish meaning:
GloVe demands that differences of vectors encode such log-ratios. That leads to the objective: dot product ≈ logarithm of the count, with a weighting f that damps the contribution of very rare and very frequent pairs:
The log-bilinear form is exactly the reason differences of embeddings correspond to ratios of probabilities — and hence why analogy arithmetic works (as in word2vec, but here derived from global statistics).
NumPy The GloVe objective
import numpy as np
# train W so that wᵢ·wⱼ + bᵢ + bⱼ ≈ log X_ij
def glove_loss(W, b, X, xmax=100, alpha=0.75):
f = np.minimum((X / xmax) ** alpha, 1.0) # weighting function
pred = W @ W.T + b[:, None] + b[None, :]
return (f * (pred - np.log(X + 1)) ** 2).sum() # weighted least squares
Why it matters
It gave a single view of count-based and predictive embeddings and explained why analogy arithmetic works (through the log-ratios). GloVe vectors were a popular pretrained input for NLP alongside word2vec, until contextual models arrived.
Connections
The same result — embeddings with linear structure — but from a different philosophy: word2vec predicts from local windows, GloVe approximates global co-occurrence statistics. It was later shown that the two are mathematically close (skip-gram implicitly factorizes a similar matrix).
The general idea "a word is a dense vector encoding meaning from context" was laid down by the neural language model. GloVe is another way of getting such vectors, this time out of global counts rather than out of training an LM.
Static embeddings (GloVe, word2vec) were the standard input layer of NLP until they were replaced by contextual representations from transformers, where a word's vector depends on the whole sentence. GloVe is the ancestor of that input layer.
Questions worth asking
If word2vec and GloVe give a similar result, what is the practical difference?
GloVe works from a global co-occurrence matrix computed in advance — that lets you parallelize and reuse the statistics, but it needs memory for the matrix. Word2vec walks the corpus window by window, streaming and frugal with memory. In practice the choice often came down to which pretrained vectors were available; quality is comparable.
Why the weighting function f(X) — can't you just minimize the error over all pairs?
Without it, very frequent pairs ("the … of" and friends) would dominate the sum, while very rare ones would add noise. f grows for moderate frequencies and saturates for huge ones — balancing the contributions. An engineering detail, but without it the quality of the embeddings drops noticeably.
Do these embeddings inherit the bias in the text?
Yes, and it is a serious problem. GloVe vectors are known to reproduce social stereotypes from the corpus (gendered associations of professions, for instance), because they simply reflect the statistics of usage. That spawned a whole line of work on measuring and reducing bias in embeddings — a representation inherits the prejudices of the data it was trained on.
What to read in the original
Read the key part — the derivation of the objective from co-occurrence ratios (the math box). It is a model example of how a simple statistical observation is turned into a trainable model.