Epilogue · 1943 → 2025

Where all this is heading

The threads running through the canon, and an honest look at where the frontier stands now.
In short. 59 papers are not 59 separate ideas but a handful of long threads that run across the decades and braid into each other. Below are those threads, and where they appear to lead next. No promises that "AI will change everything" — only what is visible from the canon itself.

The threads that run through the canon

An additive path rescues the gradient. The same "+1 in the derivative" surfaces three times: the constant error carousel in the LSTM (ct = ct−1 + …, #11), the identity shortcut in ResNet (y = x + F(x), #27) and the residual wrappers in the Transformer (#32). Depth and length were tamed not by a clever activation but by a direct path for the gradient.

Scale — and its limits. From "bigger = better" (#37 Kaplan) to "grow the data alongside the parameters" (#42 Chinchilla), and on into the wall of finite data, which opened up a new axis: spending compute at inference (#58 o1 / test-time). Every time one scaling axis hits its ceiling, the next one turns up.

Alignment keeps getting simpler. RLHF on PPO (#44, #33) → DPO drops the RL (#51) → Constitutional AI drops the human from the labeling (#45) → GRPO drops the critic (#59). The complicated three-stage pipeline is being taken apart piece by piece into cheaper, steadier parts.

The KV cache is the inference bottleneck. A whole branch of Era 6 fights over the same thing: FlashAttention makes computing attention cheaper (#48), vLLM manages the cache (#49), GQA cuts the number of KV heads (#56), MLA compresses each one (#53), prefix caching and CacheBlend reuse the cache across requests (#54). Progress here is classic systems engineering, not new architectures.

An idea waits for its substrate. LeNet's convolutions in '98 (#9) waited for GPUs and data until AlexNet in '12 (#16); the sparse MoE of 2017 (#31) ripened into Mixtral and DeepSeek (#52, #53). The right idea often arrives a decade ahead of the hardware that lets it fire.

Reasoning as a resource. Chain-of-Thought showed that reasoning step by step helps (#43); o1 turned its length into a quality dial and taught it by RL (#58); R1 and GRPO made the recipe open (#53, #59). "Think longer" went from a prompting trick to a learnable, scalable ability.

Reliability lives in the decoding layer. Structured output (#55) is a reminder: part of what makes LLM systems fit for production is born not in the model's weights but in the wrapper around it — the logit mask, the automaton, the verifier.

Where things seem to be heading

Extrapolating from the canon is a thankless business (half the papers here caught their own contemporaries by surprise). But a few directions are visible straight from its last pages:

Test-time compute and reasoning — the newest axis (#58, #59): probably a further shift of "intelligence" out of training and into inference, search and verification, and a fight to make that thinking cheap and dependable.

Agents and tool use — a direct continuation of structured output (#55) and function calling: the LLM as a component of systems that call tools and act over many steps; the bottleneck here is reliability and evaluation, not "intelligence" per se.

Inference efficiency — the KV-cache branch (#48–#57) is not closed: long context, cheap serving and low latency remain an arena for systems engineering.

The data shortage (#42) pushes towards synthetic data, training on the model's own reasoning (as in R1), and multimodality as a new source of signal.

Which of these turns out to be the next "paper #60" — time will tell. A canon is a canon precisely because it keeps being written.

Where to go from here. If you went through in order — you have come from the first formal neuron (1943) to open reasoning models (2025). If you jumped around by interest — go back to the threads above and follow one all the way through: seeing how "the additive path" or "the fight over the KV cache" runs across the eras is more useful than any single paper.