Era 1 · Origins · 1943

1 The McCulloch–Pitts neuron

A Logical Calculus of the Ideas Immanent in Nervous Activity · McCulloch & Pitts · Bull. Math. Biophysics
🟦 this write-up is enough~30–45 minoriginal ↗
The gist in 20 seconds. The first mathematical model of a neuron: a binary threshold element. A network of such elements can compute any Boolean function, and with feedback loops it works as a finite-state machine. This is the formal foundation of the "the brain is a computer" thesis. But there is no learning here — the threshold and the wiring are set by hand.

Context

1943, the height of cybernetics taking shape. Warren McCulloch (a neurophysiologist) and Walter Pitts (a self-taught logician) ask whether the nervous system can be described in the language of formal logic. In the air: Russell and Whitehead's Principia Mathematica (everything reduces to logic) and the young theory of computability (Turing, 1936). Their move: abstract the biological neuron down to the simplest possible logical device and see what can be built out of it.

The idea and the mechanism

A neuron receives binary inputs xi ∈ {0, 1} of two kinds: excitatory and inhibitory. The key point (and the essence of the model): there are no tunable weights — all excitatory inputs count equally. The neuron fires on an all-or-nothing basis if the number of active excitatory inputs reaches the threshold θ and no inhibitory input is active (any active inhibitory input is an absolute veto):

y = 1  if  (Σi xiexc ≥ θ)  and  (Σj xjinh = 0)   (otherwise y = 0)

No learning takes place: the only things "programmed" are the threshold θ and the topology of the connections (which inputs are excitatory, which inhibitory). That is already enough for a single neuron to implement the basic logic gates, and out of those, anything:

Since any Boolean function can be built from {AND, OR, NOT}, a network of MP neurons is functionally complete. And if cycles are allowed (the output fed back to an input through a delay), the network gains memory and becomes equivalent to a finite-state machine. In essence, this is computability theory in neural form.

x₁ x₂ x₃ exc.exc.inh. ⊘ Σ exc.count threshold≥ θ ? y{0,1}
The MP neuron: the number of active excitatory inputs is compared against the threshold θ (there are no weights). Any active inhibitory input (⊘) is an absolute veto — it silences the neuron no matter how much excitation arrives.
Analogy. The neuron is a turnstile with an adjustable counter: it lets you through (outputs 1) only once enough "tokens" have arrived from the inputs. The inhibitory input is the emergency stop — while it is pressed the turnstile stays locked, however many tokens you feed it.

Why it matters

This is the intellectual origin of all connectionism. The paper showed that a network of simple threshold units has universal computational power — thinking can in principle be described as computation in a network of neurons. The idea directly influenced John von Neumann (he cited it while designing the computer architecture) and cybernetics as a whole.

But the same paper carries the limitation that will define the next 15 years: the model is static. The wiring and the threshold have to be set by hand for each function. The question "how could a network tune its own weights from data?" stays open — and will be attacked by Hebb's rule, the perceptron and, finally, backpropagation.

Connections

The MP neuron can compute, but its weights are fixed — a human sets them. Hebb's rule (1949) adds exactly what is missing: a mechanism by which connections change on their own, out of the joint activity of neurons. That is the first step from a static logic circuit to a learning system — without it, the neuron remains a clever logic gate.

→ leads to3. Perceptron

The perceptron (1958) takes the same threshold neuron and equips it with a supervised learning algorithm: weights are adjusted against labeled examples, and provable convergence appears for the first time. The direct lineage 1943 → 1958 → the modern artificial neuron (linear sum + nonlinearity + weight update) starts here.

Here universality is proved for Boolean functions: a network of threshold neurons computes any logical function. Decades later the universal approximation theorem repeats the same thought for continuous functions and networks with smooth activations. Two faces of one fundamental fact: neural networks are expressively rich — the only questions are how to train them and how many neurons it takes.

What to read in the original

The original is dense formal logic in 1940s notation; reading it is optional. If you are curious, look at exactly how the authors construct logic functions out of neurons — the rest is for historians. This write-up is enough to follow the canon.