Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning technique that adapts a pre-trained neural network—most commonly a transformer—by learning a pair of low-rank matrices that approximate the required weight updates, while keeping the original model weights frozen.
What is Low-Rank Adaptation (LoRA)?
LoRA modifies a model by injecting small trainable “adapter” matrices into selected linear layers (for example, the query/key/value projections in attention or the MLP projections). Instead of updating a full weight matrix (W), LoRA represents the update as ΔW = BA, where A∈R^{r×d} and B∈R^{d×r} with rank r≪d. During fine-tuning, only A and B are trained, which dramatically reduces the number of trainable parameters and optimizer state.
A key practical benefit is that the base model remains unchanged, enabling teams to store and ship small LoRA “deltas” per task or customer. At inference time, LoRA can be merged into the base weights (for single-adapter serving) or applied on-the-fly (for dynamic adapter switching). LoRA is widely used because it often achieves performance close to full fine-tuning while lowering GPU memory requirements and training time.
Where LoRA is used and why it matters
LoRA is used to specialize large language models for domain tasks (customer support, legal drafting, medical summarization), align models with organizational style, or adapt multimodal models for specific datasets. It matters because it reduces fine-tuning cost, enables rapid iteration, and makes it feasible to maintain many task-specific variants without duplicating full model checkpoints.
Examples
- Instruction tuning a general LLM into a helpdesk assistant using LoRA on attention projection layers.
- Maintaining multiple adapters for different clients and selecting the correct adapter at runtime.
- Merging a single LoRA adapter into the base model for simplified deployment.
FAQs
Is LoRA the same as adapters? LoRA is a type of adapter method, but it specifically parameterizes updates as low-rank matrices applied to existing linear layers.
Does LoRA change the base model weights? During training, the base weights are frozen; LoRA learns separate matrices. You may optionally merge LoRA into the base weights for inference.
How do you choose the rank (r)? Higher rank increases capacity and cost. In practice, you tune rank along with which layers receive LoRA to balance quality and efficiency.
When is full fine-tuning better? If you have abundant compute, large task data, or need maximum adaptation across many layers, full fine-tuning can outperform LoRA, but it is more expensive.