Era 5 · The LLM era · 2021

41 LoRA

LoRA: Low-Rank Adaptation of Large Language Models · Hu, Shen, Wallis, Allen-Zhu et al. · Microsoft · ICLR 2022
🟧 read selectively~1 horiginal ↗
The gist in 20 seconds. Freeze the weights and insert a trainable LOW-RANK update into each layer: ~10000× fewer trainable parameters with no added inference latency. A cheap, modular fine-tune that democratized the tuning of large models.

Context

Fine-tuning a large model in full is expensive in memory and storage — a separate copy of all the weights for every task. Hu et al. make fine-tuning cheap.

The idea and the mechanism

Freeze the weights W and, to adapt, add a low-rank update ΔW = BA, where A, B are small matrices of rank r ≪ d. Only A, B are trained. The rationale: the useful task-specific weight update has a low "intrinsic rank".

linear algebra The low-rank update and the parameter saving

A full fine-tune updates the whole matrix W ∈ ℝd×k (dk parameters). LoRA freezes W and parameterises the update as the product of two "thin" matrices:

h = Wx + αr BAx,   B ∈ ℝd×r, A ∈ ℝr×k

That is r(d+k) trainable parameters instead of dk. Take d=k=4096, r=8:

8·8192 ≈ 65k  vs  4096² ≈ 16.7M  (~256× fewer)

Zero inference latency. Unlike adapter layers, BA can be folded into the weights: W′ = W + BA — and the model then runs on an ordinary matrix, with no extra operations. On top of that, many small LoRA adapters can live on one base model and be swapped per task.

PyTorch A LoRA layer
import torch, torch.nn as nn

class LoRALinear(nn.Module):
    def __init__(self, W, r=8, alpha=16):
        super().__init__()
        self.W = W                              # frozen
        d, k = W.shape
        self.A = nn.Parameter(torch.randn(r, k) * 0.01)
        self.B = nn.Parameter(torch.zeros(d, r))   # starts at zero → ΔW=0
        self.s = alpha / r
    def forward(self, x):
        return x @ self.W.T + (x @ self.A.T) @ self.B.T * self.s
x W (frozen) A B + h only the small A, B are trained (rank r)
The large matrix W is frozen; only the cheap low-rank addition BA is trained. At deployment it is folded into W — no extra latency at all.
Analogy. Rather than rewriting a whole encyclopedia for a new subject, you slip a thin errata insert into it. The encyclopedia (W) stays as it is; each task gets its own light insert (A, B). To switch subject you swap the insert instead of reprinting the volume. And when it comes to "reading", the insert can be glued in for good — turning the pages takes no longer.

Why it matters

The default method for parameter-efficient fine-tuning (PEFT); it democratized the tuning of large models — genuinely doable on a single GPU. Thousands of specialized LoRA adapters on one base LLM is now the standard shape of customization.

Connections

← applied to38. GPT-3

It was giants like GPT-3 that made a full fine-tune impractical — hence LoRA. On a 175B model LoRA cuts trainable parameters by ~10000× and memory several times over, with no loss in quality.

Two ways to make big models cheap: distillation compresses the model into a small one, LoRA adapts the large one cheaply without touching its weights. They are often combined: a distilled base is then tuned with LoRA for specific tasks.

→ a tool for50. LLaMA

Open weights of LLaMA + LoRA = the explosion of custom models: the community tunes the base LLaMA for thousands of tasks with light adapters on consumer hardware. LoRA is a key reason the open ecosystem was able to iterate so fast.

Questions worth asking

Why should a task-specific weight update be low-rank in the first place?

An empirical hypothesis from the paper: adaptation to a particular task "lives" in a small subspace — a large pre-trained model already knows nearly everything, and fine-tuning only nudges it into position. The authors showed that even rank 1–2 often works. There is no rigorous theory, but the low-rankness of updates holds up in practice across many tasks.

Why is B initialized to zero and A at random?

So that at the start of training ΔW = BA = 0 and the model begins from exactly the pre-trained behavior, without a jolt. A is random (otherwise the gradient with respect to B would be zero — a symmetry), B is zero (so the start is clean). Fine-tuning thus begins at a known good point and moves away from it smoothly.

If BA folds into the weights with no latency, where is the catch?

The catch is flexibility at inference. Once you have folded an adapter in, the model is specialized — to serve many tasks at once (multi-LoRA in a single service) the adapters have to be applied dynamically, and then a small overhead does appear. "Zero latency" is true for one fixed adapter; multiplexing is an engineering problem of its own.

What to read in the original

Read the essentials — the low-rank idea and why there is no latency; the experiments on choosing the rank r are useful for intuition.