50 LLaMA
Context
The strong LLMs were closed (GPT-3/3.5). Meta releases a family of efficient foundation models trained on public data.
The idea and the mechanism
Following the lessons of Chinchilla (#42), they take smaller models (7B–65B) but train them on MORE tokens → a better deal in terms of INFERENCE compute. The 13B LLaMA beats GPT-3 175B on many benchmarks. The architectural details became the standard: RMSNorm (normalization), SwiGLU (the FFN activation), RoPE (rotary positional embeddings).
linear algebra RoPE: position as a rotation of vectors
How do you encode position? RoPE rotates pairs of coordinates in Q and K by an angle proportional to the position. For a pair (x1, x2) at position m:
The magic is that the dot product of the rotated qm and kn depends only on the relative position (m−n), not on the absolute ones. That gives you relative positioning "for free" (no separate embeddings) and extrapolates better to lengths not seen in training. Add RMSNorm (normalizing by the root mean square — cheaper than LayerNorm) and SwiGLU (a gated activation, empirically better than ReLU) — these three choices became the de facto standard architecture for open LLMs.
PyTorch RoPE — rotating Q/K by position
import torch
# rotate pairs of dimensions by an angle ∝ position
def rope(x, pos, theta):
ang = pos * theta # the angle grows with position
x1, x2 = x[..., 0::2], x[..., 1::2]
return torch.cat([x1*ang.cos() - x2*ang.sin(),
x1*ang.sin() + x2*ang.cos()], dim=-1)
# after RoPE: q_m · k_n depends only on (m − n)
Why it matters
The weights (under a research licence) leaked in March 2023 → an open-source explosion: Alpaca, Vicuna and thousands of fine-tunes; it started the wave of open LLMs that Mistral and the rest grew out of. LLaMA-2 and 3 are officially open.
Connections
LLaMA is Chinchilla's lesson made flesh: smaller models on far more tokens. And it goes further — it trains past the Chinchilla point for the sake of cheap inference, which pays for itself over millions of requests.
Open LLaMA weights + LoRA = an explosion of customization: the community fine-tunes the base with lightweight adapters on consumer hardware. An open model and cheap fine-tuning together are what made iteration this fast.
An open model has to be run somewhere efficiently — vLLM became the standard serving engine for the LLaMA family. Open weights + efficient inference = practical self-hosting.
Questions worth asking
Why was the "weights leak" a turning point rather than just an incident?
Because for the first time the community had a strong base model it could freely fine-tune and run. Before that, open models lagged noticeably behind. LLaMA gave a foundation on which Alpaca, Vicuna and hundreds of projects grew within weeks — which shifted the center of innovation partly out of the labs and into open source. Meta later drew its conclusions and released LLaMA-2 and 3 openly by design.
Why train past the Chinchilla point if it is "optimal"?
Chinchilla optimizes training under a fixed compute budget. But for a deployed model what matters more is the cost of inference (millions of requests). A small model is cheaper to serve, so it pays to "over-train" it with extra tokens: you spend more on training once and save forever on requests. LLaMA deliberately goes against compute-optimal in favor of deploy-optimal.
RMSNorm, SwiGLU, RoPE — are these fundamental improvements or minor hacks?
More like carefully selected empirical improvements: each gives a small gain, but together they are noticeable and cost almost no extra compute. RMSNorm is cheaper, SwiGLU is slightly better than ReLU, RoPE is better for long contexts. Their value is that they became a standard set: almost every modern open LLM uses exactly this trio.
What to read in the original
Read the essentials — the data and architecture choices (RMSNorm, SwiGLU, RoPE) and the "smaller, but trained longer" results.