Skip to main content

📝 Low-Rank Adaptation (LoRA)

Description​

< What is it? >​

Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning method. It freezes a pretrained weight matrix and learns a small, low-rank update instead of training a second full-sized matrix.

This makes fine-tuning use much less trainable memory and produces a small adapter checkpoint that can be loaded alongside the same base model.

Key points​

< The low-rank update >​

For a base matrix W0∈Rdout×dinW_0\in\mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}}, LoRA uses an effective weight:

W=W0+ΔW=W0+αrBAW = W_0+\Delta W = W_0+\frac{\alpha}{r}BA

where A∈Rr×dinA\in\mathbb{R}^{r\times d_{\mathrm{in}}}, B∈Rdout×rB\in\mathbb{R}^{d_{\mathrm{out}}\times r}, and the rank rr is much smaller than dind_{\mathrm{in}} and doutd_{\mathrm{out}}. The orientation may be transposed in codebases that use row-vector notation, but the idea is unchanged: two small matrices produce the update.

  • Example: if W0W_0 is 4096×40964096\times4096 and r=8r=8, a full update needs 16,777,21616{,}777{,}216 parameters, while the LoRA update needs only 8(4096+4096)=65,5368(4096+4096)=65{,}536 parameters, about 0.39%0.39\% as many.

< What trains and what stays frozen? >​

W0W_0 stays frozen; only AA and BB receive gradients. A common initialization makes ΔW=0\Delta W=0 at the start of training, so the model initially behaves exactly like the base checkpoint.

LoRA is often applied to selected Transformer projection matrices, such as query and value projections. The target modules are a hyperparameter: applying LoRA to more modules gives more adaptation capacity but increases the adapter size and memory use.

< Rank and scaling >​

  • Rank rr controls adapter capacity. A larger rank can learn a larger change but costs more memory and can overfit more easily.
  • Scaling α\alpha controls the effective size of the update through α/r\alpha/r.
  • LoRA dropout optionally regularizes the adapter input during training.

Tune rank, scaling, target modules, learning rate, and dropout on validation data rather than copying one recipe to every model and task.

< Training and serving >​

During training, frozen base weights need no gradients or optimizer state, which greatly reduces memory use. The adapter checkpoint contains only the learned LoRA parameters, so many task-specific adapters can share one base model.

At inference, an adapter can remain separate for quick switching or be merged into the base weight matrix. Merging removes adapter-selection flexibility but can simplify deployment.

< LoRA and QLoRA >​

QLoRA is LoRA with a quantized frozen base model, commonly 4-bit during training. The LoRA adapters still train in a higher-precision representation. It reduces base-model memory further, but quantization introduces hardware, kernel, and numerical constraints.

< Limitations >​

LoRA is not automatically equivalent to full fine-tuning. A very low rank or poorly chosen target modules may not express the required behavior change. It also does not fix weak data, evaluation leakage, harmful behavior, or an unsuitable base model.

Comparison​

< LoRA compared with full fine-tuning and QLoRA >​

MethodBase weightsTrainable parametersMain tradeoff
Full fine-tuningUpdatedAll or most parametersMaximum flexibility, highest memory and checkpoint cost
LoRAFrozenLow-rank adapter matricesEfficient and swappable, but capacity depends on rank and target modules
QLoRAQuantized and frozenLow-rank adapter matricesLowest memory, with more quantization constraints

Reference​