LoRA trains small adapter layers so you customize a model without touching its full weights.
DefinitionWhat it means
LoRA, Low-Rank Adaptation, is the most widely used form of PEFT, Parameter-Efficient Fine-Tuning. Instead of updating every weight in a large model, it freezes the original weights and trains small low-rank matrices inserted alongside them, often under one percent of the original parameter count. At inference time the adapter can be merged in or kept separate and swapped out.
Why it mattersWhy you should care
LoRA makes fine-tuning affordable and fast enough to run on modest hardware, and because adapters are small files, teams can maintain many task-specific or customer-specific adapters on top of one shared base model rather than hosting dozens of full copies. This is the default approach for customizing open-weight models in production today.
At a glanceSee it
Inside a layer, LoRA leaves W frozen and learns only the tiny rank-r detour BA — so the update h equals Wx plus a scaled BAx.
At serving time the real choice is hot-swapping many adapters over one shared base versus merging one back into W for latency-free inference.
Where you see itIn the wild
- Open-weight model fine-tuning jobs on a single GPU or a small cluster.
- Multi-tenant platforms swapping customer-specific adapters over one base model.
- Hugging Face PEFT library configs specifying rank and target modules.