4 Perceptrons (the critique)
Context
By the late 1960s there is a lot of noise around perceptrons and very little rigour. Marvin Minsky and Seymour Papert (MIT) run a cold mathematical analysis: what can these machines do, and what is out of reach in principle. The book is dedicated to... Rosenblatt himself: they had been schoolmates and friendly rivals.
The idea and the mechanism
A perceptron draws one separating line. That is enough for AND and OR, but not for XOR (exclusive OR), where the members of each class sit on opposite corners of a square and interleave along the diagonals. More generally, Minsky and Papert showed that predicates such as parity and connectedness of a figure are unreachable for a perceptron with a local receptive field.
logic · algebra Proof: XOR cannot be realised by a single perceptron
Suppose a perceptron with weights w1, w2 and threshold θ realises XOR. The four rows of the truth table give four inequalities:
Add the second and the third: w1 + w2 ≥ 2θ. But the fourth demands w1 + w2 < θ. So 2θ ≤ w1+w2 < θ, hence 2θ < θ, that is θ < 0 — a contradiction with the first inequality θ > 0. ∎
No straight line separates the XOR points. A hidden layer (see the code below), on the other hand, builds two intermediate features — and the problem disappears.
NumPy How two layers solve XOR
import numpy as np
step = lambda z: (z >= 0).astype(float)
X = np.array([[0,0],[0,1],[1,0],[1,1]])
# the hidden layer builds two features: h1 = OR, h2 = AND
W1 = np.array([[1,1],[1,1]]); b1 = np.array([-0.5, -1.5])
# output: XOR = OR and NOT AND → h1 − h2 > 0
W2 = np.array([1,-1]); b2 = -0.5
h = step(X @ W1.T + b1) # [OR, AND] for each point
y = step(h @ W2 + b2) # XOR
print(y) # → [0, 1, 1, 0] ✓
Why it matters
The book (plus the fact that nobody could train multi-layer networks then) undermined confidence and funding — the first AI winter set in for roughly 15 years, and attention moved to symbolic AI. A lesson for all time: authoritative criticism of a correct direction can hold it back for a decade — and a limitation that looked fundamental is lifted by a single hidden layer.
Connections
It is the perceptron's capabilities that get dissected to the limit here. The perceptron convergence theorem works on linearly separable data — Minsky and Papert mark the boundary of that class with XOR.
A hidden layer solves XOR (see the code), but you have to be able to train it. The moment backprop taught multi-layer networks to learn, Minsky and Papert's pessimism collapsed and the winter ended.
One of the works that brought interest back to neural networks in the 1980s — coming from physics, going around the "perceptron" agenda the criticism had focused on.
Questions worth asking
If two layers solve XOR (the code is right there), why did the criticism do so much damage?
Because the problem was not representation but learning. Setting the hidden layer's weights by hand is easy — but in 1969 there was no method to teach them from data (backprop was not yet known or popular). A perceptron could learn; a multi-layer network could not. Add the rigour of the book, the authority of its authors and the politics of funding. The gap was precisely in the algorithm for training hidden layers.
Did Minsky and Papert really "kill" neural networks — or is that a simplification?
A simplification. The downturn had more than one source: there was also the Lighthill report, a general disillusionment, and grant cuts. The book rigorously handled the single-layer case and merely conjectured (wrongly) that the multi-layer case was barren too. That is speculation, not a theorem — but in the atmosphere of 1969 it landed as a verdict.
Are there modern "XORs" — problems where a simple model hits the same wall?
Yes, and they are instructive. Any non-linearly-separable structure is XOR writ large: the parity of n bits, for instance, is extremely hard for many models without explicit nonlinearity. The modern echo is the limited reach of linear probes on top of embeddings, and the recent "grokking" phenomenon on modular arithmetic, where a network fails to "get" the task for a long time and then suddenly generalizes. The XOR lesson — linear is not enough — has not gone anywhere.
What to read in the original
The gist plus the XOR proof (the math box) is enough. There is no need to go deep into the geometry of predicates (order, diameter) unless you are a historian or a theorist. The main thing is to understand both the shape of the limitation and the inflated reading of it.