41 LoRA
Context
Fine-tuning a large model in full is expensive in memory and storage — a separate copy of all the weights for every task. Hu et al. make fine-tuning cheap.
The idea and the mechanism
Freeze the weights W and, to adapt, add a low-rank update ΔW = BA, where A, B are small matrices of rank r ≪ d. Only A, B are trained. The rationale: the useful task-specific weight update has a low "intrinsic rank".
linear algebra The low-rank update and the parameter saving
A full fine-tune updates the whole matrix W ∈ ℝd×k (dk parameters). LoRA freezes W and parameterises the update as the product of two "thin" matrices:
That is r(d+k) trainable parameters instead of dk. Take d=k=4096, r=8:
Zero inference latency. Unlike adapter layers, BA can be folded into the weights: W′ = W + BA — and the model then runs on an ordinary matrix, with no extra operations. On top of that, many small LoRA adapters can live on one base model and be swapped per task.
PyTorch A LoRA layer
import torch, torch.nn as nn
class LoRALinear(nn.Module):
def __init__(self, W, r=8, alpha=16):
super().__init__()
self.W = W # frozen
d, k = W.shape
self.A = nn.Parameter(torch.randn(r, k) * 0.01)
self.B = nn.Parameter(torch.zeros(d, r)) # starts at zero → ΔW=0
self.s = alpha / r
def forward(self, x):
return x @ self.W.T + (x @ self.A.T) @ self.B.T * self.s
Why it matters
The default method for parameter-efficient fine-tuning (PEFT); it democratized the tuning of large models — genuinely doable on a single GPU. Thousands of specialized LoRA adapters on one base LLM is now the standard shape of customization.
Connections
It was giants like GPT-3 that made a full fine-tune impractical — hence LoRA. On a 175B model LoRA cuts trainable parameters by ~10000× and memory several times over, with no loss in quality.
Two ways to make big models cheap: distillation compresses the model into a small one, LoRA adapts the large one cheaply without touching its weights. They are often combined: a distilled base is then tuned with LoRA for specific tasks.
Open weights of LLaMA + LoRA = the explosion of custom models: the community tunes the base LLaMA for thousands of tasks with light adapters on consumer hardware. LoRA is a key reason the open ecosystem was able to iterate so fast.
Questions worth asking
Why should a task-specific weight update be low-rank in the first place?
An empirical hypothesis from the paper: adaptation to a particular task "lives" in a small subspace — a large pre-trained model already knows nearly everything, and fine-tuning only nudges it into position. The authors showed that even rank 1–2 often works. There is no rigorous theory, but the low-rankness of updates holds up in practice across many tasks.
Why is B initialized to zero and A at random?
So that at the start of training ΔW = BA = 0 and the model begins from exactly the pre-trained behavior, without a jolt. A is random (otherwise the gradient with respect to B would be zero — a symmetry), B is zero (so the start is clean). Fine-tuning thus begins at a known good point and moves away from it smoothly.
If BA folds into the weights with no latency, where is the catch?
The catch is flexibility at inference. Once you have folded an adapter in, the model is specialized — to serve many tasks at once (multi-LoRA in a single service) the adapters have to be applied dynamically, and then a small overhead does appear. "Zero latency" is true for one fixed adapter; multiplexing is an engineering problem of its own.
What to read in the original
Read the essentials — the low-rank idea and why there is no latency; the experiments on choosing the rank r are useful for intuition.