Era 6 · Generative models and systems · 2024 / 2025

53 DeepSeek V3 / R1

DeepSeek-V3 Technical Report (Dec 2024) · DeepSeek-R1 (Jan 2025) · DeepSeek-AI
🟧 read selectively~2–3 horiginal ↗
The gist in 20 seconds. V3 — a 671B MoE (37B active) with Multi-head Latent Attention and expert load balancing without an auxiliary loss; a strong open-weight model trained on (reportedly) modest compute. R1 — reasoning grown out of PURE RL on verifiable rewards: self-checking emerges on its own. OpenAI o1 level, open weights.

Context

By 2024–25 the frontier was mostly closed (GPT-4, Claude, Gemini). DeepSeek ships open models of comparable quality and loudly advertises how little the training cost.

The idea and the mechanism

V3 is the base: a 671B-parameter MoE (~37B active per token) with two engineering twists — Multi-head Latent Attention (MLA) and expert load balancing without an auxiliary loss. R1 is the reasoner: strong reasoning is grown by almost pure RL on verifiable rewards (RLVR), all the way to the R1-Zero variant with no prior SFT at all.

reinforcement learning MLA (compressing the KV) and RLVR (reasoning out of RL)

MLA — Multi-head Latent Attention. Instead of caching the full K, V for every head, they are compressed into a shared low-rank latent c (down-projection); only c is cached, and at attention time it is expanded back out (up-projection):

c = Wdown h  (cached),   K, V = Wup c  (reconstructed)

The KV cache shrinks by a large factor → long context gets cheaper (compare PagedAttention #49, except here the saving is in the architecture itself).

RLVR — RL on Verifiable Rewards. The reward is the checkable correctness of the final answer (math, code — anywhere the result can be verified), not the score of a reward model:

r(answer) = 1 if the answer is correct (verifiably); otherwise 0

R1-Zero was trained by PURE RL from the base model, with no SFT — and reasoning behaviors (self-checking, reconsidering, lengthening the chain) EMERGED on their own, including the famous "aha moment". This echoes AlphaGo (#29): improvement through interaction + verification, only now the "search" has been unrolled into a chain of thought.

Python RLVR — reward for verifiable correctness
# no reward model: the reward is whether the answer is right (checkable)
def reward(answer, question):
    return 1.0 if verify(answer, question.gold) else 0.0   # math/code
# R1-Zero: pure RL from the base model, no SFT
# → reasoning (self-checking, long chains) emerges by itself
base (MoE+MLA)V3 pure RL (RLVR)verifiable rewards reasoningR1 ("aha") reasoning grew out of RL, not out of labels
A strong open base (V3) + pure RL on verifiable rewards (R1) → the ability to reason emerges on its own, with no SFT demonstrations.
Analogy. Rather than drilling a student on model solutions (SFT), just hand them problems with a checkable answer and reward correctness. To hit the answer more often, the student STARTS writing out the working, double-checking, going back over mistakes — the reasoning strategy is born out of wanting the right answer, not copied off examples. Same spirit as self-play in AlphaGo.

Why it matters

It pulled open models up close to the frontier; R1 is a public recipe for reasoning via RL (a direct continuation of Chain-of-Thought #43 plus the "learning + search" idea from AlphaGo). One caveat (flag): the quoted cost of V3's final run (~2.788M H800-hours ≈ $5.6M) is DISPUTED — it excludes the prior research, the failed runs, the data and the capex; it is the cost of the final run, not the all-in total.

Connections

V3 is the sparse MoE of "Outrageously Large" (2017) carried to the frontier: 671B total, ~37B active, plus routing improvements (load balancing without an auxiliary loss). A straight line from the idea of conditional computation to an open GPT-4-class model.

CoT (2022) showed that reasoning step by step lifts quality sharply — but through the prompt. R1 takes the next step: it trains the model to reason via RL, and long chains of thought appear on their own. Prompting turned into a learned ability.

↔ echoes29. AlphaGo

The same philosophy of "improvement through interaction + verification": AlphaGo is self-play + tree search, R1 is RL + verifiable rewards, where the "search" is unrolled into a chain of reasoning. Nine years on, the idea came back from Go to language.

Questions worth asking

Why does reasoning "emerge by itself" from pure RL — is it really not programmed in?

No, it is not programmed in — it is emergent. The model is simply rewarded for a correct answer; to hit it more often on hard problems, spending more "thought" pays off — writing out steps, checking itself, trying alternatives. RL finds that strategy on its own, with no explicit instruction. What is striking is that self-checking and longer reasoning arise as an instrumental goal of maximizing reward.

Why does RLVR work where ordinary RLHF stalls?

Because the reward is objective and checkable (the answer is right or it isn't), not the score of an imperfect reward model. There is no reward hacking in the usual sense — you cannot "fool" a correctness check. The limitation: RLVR only applies where the result is verifiable (math, code, logic); on "soft" tasks (style, helpfulness) you still need human or AI preferences.

Did DeepSeek really cost ~$6M — and why did that crash the market?

The ~$5.6M figure is the cost of the final run at rental rates, excluding the prior research, the failed runs, the data and the capital cost of the cluster. The full cost is many times higher. But the mere fact that a frontier-class open model had been trained cheaply sent stocks tumbling in January 2025 (Nvidia −~17% in a day, a record loss of market cap) — the market took fright that the monopoly of expensive closed labs, and demand for GPUs, were under threat. The lesson: the engineering number matters and so does its interpretation — and here the interpretation ran well ahead of the number.

What to read in the original

V3 — skim it (MLA, the MoE engineering); R1 — read the key parts (the pure-RL reasoning setup, RLVR and the description of the "aha moment"). This is where the canon ends — and a good place from which today's frontier begins.