Era 1 · Origins · 1982

6 Hopfield network

Neural Networks and Physical Systems with Emergent Collective Computational Abilities · John Hopfield · PNAS
🟧 read selectively~1–1.5 horiginal ↗
The gist in 20 seconds. A physicist came up with a network that has an "energy" and always rolls down into a minimum. If the minima are stored patterns, you get associative memory: feed in a corrupted image → the network reconstructs the nearest intact one. A bridge between neural networks and statistical physics, which won a Nobel Prize in physics in 2024.

Context

After the winter neural networks are out of favor, but John Hopfield — a physicist — looks at them through the eyes of statistical mechanics and triggers the 1980s renaissance. His move: introduce a quantity called "energy" and show that the network's dynamics minimize it.

The idea and the mechanism

A network of N binary neurons si ∈ {−1, +1} with symmetric weights (wij = wji, wii = 0). The energy is defined as:

E = −12 Σi,j wij si sj

Asynchronous updates si ← sign(Σj wij sj) can only lower E (see the math box), so the network is guaranteed to converge to a local minimum. The weights are set by the Hebb rule for a set of patterns — and those patterns become minima, that is attractors. This gives you content-addressable memory: a corrupted input rolls down to the nearest stored image.

statistical physics Why the energy cannot go up (a Lyapunov function)

Suppose neuron i is updated: si → si′. Introduce the local field hi = Σj≠i wij sj. Thanks to the symmetry and to wii=0, the change in energy depends only on this neuron:

ΔE = Enew − Eold = −(si′ − si) · hi = −Δsi · hi

The update rule sets si′ = sign(hi). If the neuron flipped, its sign previously disagreed with hi and now agrees — so Δsi has the same sign as hi:

Δsi · hi ≥ 0  ⟹   ΔE ≤ 0

The energy is monotonically non-increasing and bounded below — so the dynamics converge to a fixed point (a local minimum). E is a Lyapunov function of the system. ∎

NumPy Implementation: store patterns and restore a corrupted one
import numpy as np

def train(P):                        # P: (k, N) patterns of ±1
    W = sum(np.outer(p, p) for p in P) / len(P)   # Hebb rule
    np.fill_diagonal(W, 0)           # no self-connections
    return W

def recall(W, s, steps=10):
    s = s.copy()
    for _ in range(steps):
        for i in np.random.permutation(len(s)):   # asynchronously
            s[i] = 1.0 if W[i] @ s >= 0 else -1.0
    return s                          # ≈ the nearest stored pattern
energy E pattern A pattern B spurious min. corrupted input
The energy landscape. A corrupted input rolls down into the nearest valley — a stored pattern. Extra patterns create spurious minima.
Analogy. A bumpy landscape full of hollows. Drop a ball anywhere and it rolls into the nearest hollow. Each hollow is a memory; to "recall" is to roll into it from a similar state. Overload the landscape with hollows and parasitic dips appear, and the ball gets stuck in the wrong one.

Why it matters

Hopfield tied neural networks to the physics of spin glasses (the Ising model): converging to a memory = a system relaxing into a low-energy state. That legitimized neural networks among natural scientists and supplied an analytical toolkit. The stochastic extension — Hinton's Boltzmann machine — can already learn. Hopfield and Hinton received the 2024 Nobel Prize in physics for this work.

Connections

← builds on2. Hebbian learning

The network's weights are exactly the Hebb rule applied to the patterns being stored (a sum of outer products). Hebb's abstract postulate becomes a concrete mechanism for writing to memory here.

↔ part of the thaw4. Perceptrons (the critique)

One of the works that brought interest back to neural networks after the Minsky–Papert winter — but from an unexpected direction, through statistical physics rather than along the perceptron line.

→ echoes32. Transformer

"Modern Hopfield networks" (2020), with continuous states, have exponential capacity, and their update rule coincides mathematically with the Transformer's softmax attention. Attention is, in essence, a single-step retrieval from associative memory.

Questions worth asking

Why must the weights be symmetric? What breaks without that?

Symmetry is what guarantees that an energy — a Lyapunov function — exists, and hence that the network converges to fixed points. Without it the proof that ΔE ≤ 0 falls apart, and the network can cycle or behave chaotically. Sometimes that is even what you want (asymmetric networks are used to generate sequences), but then you lose the clean "memory as a minimum" picture.

Where does the capacity ≈ 0.14N come from — why so little?

Writing several patterns into the same weights creates cross-talk: the other patterns distort the local field. Statistical-physics analysis (Amit–Gutfreund–Sompolinsky) gives a critical load of ≈ 0.138·N random patterns, past which recall errors grow avalanche-like and spurious minima dominate. The memory does not "forget a bit at a time" — it breaks sharply beyond the threshold.

Is the link to attention serious mathematics or a pretty metaphor?

Serious. In "Modern Hopfield Networks" (Ramsauer et al., 2020) the states are continuous and the update rule, derived from the energy, turns out to be literally softmax(QKᵀ)V — the Transformer's attention. So attention can be read as one step of retrieval from an associative memory with exponential capacity. That is a formal correspondence, not an analogy.

What to read in the original

The paper is short (~5 pages) and elegant — the key passages are worth reading: the definition of the energy and the proof that it decreases. The best way to get a feel for why "dynamics = energy minimization" is so powerful.