📝 Low-Rank Adaptation (LoRA)
Description
< What is it? >
Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning method. It freezes a pretrained weight matrix and learns a small, low-rank update instead of training a second full-sized matrix.
This makes fine-tuning use much less trainable memory and produces a small adapter checkpoint that can be loaded alongside the same base model.
Key points
< The low-rank update >
For a base matrix , LoRA uses an effective weight:
where , , and the rank is much smaller than and . The orientation may be transposed in codebases that use row-vector notation, but the idea is unchanged: two small matrices produce the update.
- Example: if is and , a full update needs parameters, while the LoRA update needs only parameters, about as many.
< What trains and what stays frozen? >
stays frozen; only and receive gradients. A common initialization makes at the start of training, so the model initially behaves exactly like the base checkpoint.
LoRA is often applied to selected Transformer projection matrices, such as query and value projections. The target modules are a hyperparameter: applying LoRA to more modules gives more adaptation capacity but increases the adapter size and memory use.
< Rank and scaling >
- Rank controls adapter capacity. A larger rank can learn a larger change but costs more memory and can overfit more easily.
- Scaling controls the effective size of the update through .
- LoRA dropout optionally regularizes the adapter input during training.
Tune rank, scaling, target modules, learning rate, and dropout on validation data rather than copying one recipe to every model and task.
< Training and serving >
During training, frozen base weights need no gradients or optimizer state, which greatly reduces memory use. The adapter checkpoint contains only the learned LoRA parameters, so many task-specific adapters can share one base model.
At inference, an adapter can remain separate for quick switching or be merged into the base weight matrix. Merging removes adapter-selection flexibility but can simplify deployment.
< LoRA and QLoRA >
QLoRA is LoRA with a quantized frozen base model, commonly 4-bit during training. The LoRA adapters still train in a higher-precision representation. It reduces base-model memory further, but quantization introduces hardware, kernel, and numerical constraints.
< Limitations >
LoRA is not automatically equivalent to full fine-tuning. A very low rank or poorly chosen target modules may not express the required behavior change. It also does not fix weak data, evaluation leakage, harmful behavior, or an unsuitable base model.
Comparison
< LoRA compared with full fine-tuning and QLoRA >
| Method | Base weights | Trainable parameters | Main tradeoff |
|---|---|---|---|
| Full fine-tuning | Updated | All or most parameters | Maximum flexibility, highest memory and checkpoint cost |
| LoRA | Frozen | Low-rank adapter matrices | Efficient and swappable, but capacity depends on rank and target modules |
| QLoRA | Quantized and frozen | Low-rank adapter matrices | Lowest memory, with more quantization constraints |
Related ideas
- Parameter-Efficient Fine-Tuning (PEFT) compares LoRA with other efficient adaptation methods.
- Fine-Tuning explains the broader training workflow and data requirements.
- Transformer explains the projection matrices often targeted by LoRA.
- Hugging Face TRL + PEFT introduces a common training stack.