Era 3 · The deep learning explosion · 2014

24 GloVe

GloVe: Global Vectors for Word Representation · Pennington, Socher & Manning · EMNLP
🟧 read selectively~40 minoriginal ↗
The gist in 20 seconds. Embeddings from global co-occurrence statistics: vectors are trained so that their dot product approximates the logarithm of the count. The key idea — meaning is carried by the ratios of co-occurrence probabilities. It unifies the count-based and the predictive approaches.

Context

By 2014 there are two camps in embeddings: count-based methods (working off the co-occurrence matrix, like LSA) and predictive ones (word2vec). Pennington, Socher and Manning at Stanford put them together.

The idea and the mechanism

We build a global co-occurrence matrix Xij (how often word j appears in the context of i across the whole corpus) and train vectors so that their dot product approximates the logarithm of the count. Quality is on a par with word2vec, but training runs on global statistics rather than local windows.

linear algebra Why ratios of probabilities — and the log-bilinear form

Compare two words through a "probe" word k. A probability on its own, P(k|i), tells you little; their ratio, however, does distinguish meaning:

P(solid | ice)P(solid | steam) ≫ 1,   P(gas | ice)P(gas | steam) ≪ 1,   P(water | ice)P(water | steam) ≈ 1

GloVe demands that differences of vectors encode such log-ratios. That leads to the objective: dot product ≈ logarithm of the count, with a weighting f that damps the contribution of very rare and very frequent pairs:

J = Σi,j f(Xij) (wi·wj + bi + bj − log Xij)²

The log-bilinear form is exactly the reason differences of embeddings correspond to ratios of probabilities — and hence why analogy arithmetic works (as in word2vec, but here derived from global statistics).

NumPy The GloVe objective
import numpy as np
# train W so that wᵢ·wⱼ + bᵢ + bⱼ ≈ log X_ij
def glove_loss(W, b, X, xmax=100, alpha=0.75):
    f = np.minimum((X / xmax) ** alpha, 1.0)        # weighting function
    pred = W @ W.T + b[:, None] + b[None, :]
    return (f * (pred - np.log(X + 1)) ** 2).sum()  # weighted least squares
probe k:solidgaswater P(k|ice)/P(k|steam) ≫ 1≪ 1≈ 1 tells "ice" (a solid) apart……from "steam" (a gas); "water" is neutral to both
Meaning is carried not by the probabilities themselves but by their ratios: "solid" separates ice from steam sharply, "water" does not. GloVe trains vectors to reproduce those log-ratios.
Analogy. To work out the difference between two people it is useless to ask "does he enjoy food?" (both do). What helps are discriminating questions: "who is at the gym more often?". GloVe looks for exactly those probe words whose frequency ratios push concepts into different corners of the space.

Why it matters

It gave a single view of count-based and predictive embeddings and explained why analogy arithmetic works (through the log-ratios). GloVe vectors were a popular pretrained input for NLP alongside word2vec, until contextual models arrived.

Connections

↔ close relative17. Word2Vec

The same result — embeddings with linear structure — but from a different philosophy: word2vec predicts from local windows, GloVe approximates global co-occurrence statistics. It was later shown that the two are mathematically close (skip-gram implicitly factorizes a similar matrix).

← inherits from13. Neural language model

The general idea "a word is a dense vector encoding meaning from context" was laid down by the neural language model. GloVe is another way of getting such vectors, this time out of global counts rather than out of training an LM.

→ leads to32. Transformer

Static embeddings (GloVe, word2vec) were the standard input layer of NLP until they were replaced by contextual representations from transformers, where a word's vector depends on the whole sentence. GloVe is the ancestor of that input layer.

Questions worth asking

If word2vec and GloVe give a similar result, what is the practical difference?

GloVe works from a global co-occurrence matrix computed in advance — that lets you parallelize and reuse the statistics, but it needs memory for the matrix. Word2vec walks the corpus window by window, streaming and frugal with memory. In practice the choice often came down to which pretrained vectors were available; quality is comparable.

Why the weighting function f(X) — can't you just minimize the error over all pairs?

Without it, very frequent pairs ("the … of" and friends) would dominate the sum, while very rare ones would add noise. f grows for moderate frequencies and saturates for huge ones — balancing the contributions. An engineering detail, but without it the quality of the embeddings drops noticeably.

Do these embeddings inherit the bias in the text?

Yes, and it is a serious problem. GloVe vectors are known to reproduce social stereotypes from the corpus (gendered associations of professions, for instance), because they simply reflect the statistics of usage. That spawned a whole line of work on measuring and reducing bias in embeddings — a representation inherits the prejudices of the data it was trained on.

What to read in the original

Read the key part — the derivation of the objective from co-occurrence ratios (the math box). It is a model example of how a simple statistical observation is turned into a trainable model.