Home › Fine-Tuning & Alignment › PEFT / LoRA
🎯 · Models

PEFT / LoRA

Freeze the base model and train small adapter matrices instead of every weight.

In one line

Get most of fine-tuning's benefit while training a tiny fraction of the parameters.

ConceptWhat it is

PEFT, parameter-efficient fine-tuning, and its most popular method LoRA, low-rank adaptation, freeze the pretrained model's weights entirely and instead train small, low-rank matrices injected into specific layers. It exists because full fine-tuning's compute and storage cost is unnecessary for most task adaptation, and low-rank updates capture the needed behavior change with far fewer trainable parameters.

Because the base model stays frozen, multiple LoRA adapters can be trained cheaply and swapped in at serving time without duplicating the whole model.

How it worksThe mechanics

Small trainable low-rank matrices are inserted alongside the frozen weight matrices in attention or feed-forward layers; during training only these adapter matrices receive gradient updates, and at inference time the adapter's output is added to the frozen layer's output, either merged into the weights or kept separate for easy swapping.

At a glanceSee it

PEFT / LoRA diagram
PEFT / LoRA diagram 1

Inside one LoRA layer — the frozen weight update is factored into a down-projection then an up-projection, so only the skinny rank-r bottleneck ever trains.

PEFT / LoRA diagram 2

A selection path from memory budget and how many swappable task variants you need to the right point on the full-fine-tune to QLoRA spectrum.

When to use itWhere it fits

  • Adapting a large model to a narrow task or domain style without full retraining budget.
  • Teams needing to serve many task-specific variants of the same base model cheaply.
  • Limited labeled data, where a smaller parameter count reduces overfitting risk.
  • Fast iteration cycles where training runs need to complete in hours, not days.

When NOT to use itLimits & anti-patterns

  • Tasks requiring a deep shift in the model's core capabilities, where low-rank updates are not expressive enough.
  • Cases where the base model's knowledge itself is wrong or missing entirely, which retrieval or full fine-tuning handles better.
  • Extremely latency-sensitive serving where even the small overhead of adapter application matters, though this is rare in practice.

Trade-offsAdvantages & costs

Advantages
  • Trains a tiny fraction of parameters, cutting compute and memory cost dramatically.
  • Multiple adapters can be stored and swapped cheaply on top of one shared base model.
  • Much lower risk of catastrophic forgetting than full fine-tuning.
  • Fast enough to iterate on a single GPU for many practical model sizes.
Trade-offs & costs
  • Ceiling on performance can be lower than full fine-tuning for very large distribution shifts.
  • Requires choosing rank and target layers, adding a hyperparameter search.
  • Slight inference overhead if adapters are not merged into base weights.
  • Still requires a labeled, task-specific dataset to be effective.

ExampleIn the real world

Many open-weight model fine-tunes on Hugging Face, like domain-specific chat variants of Llama, are trained with QLoRA on a single consumer GPU rather than full-parameter retraining.

ToolsHow to implement it

  • Hugging Face PEFTthe standard library implementing LoRA and other parameter-efficient methods.
  • QLoRAcombines LoRA with quantization to fine-tune large models on modest hardware.
  • Axolotlpopular open-source framework for configuring and running LoRA fine-tuning jobs.
  • bitsandbytesquantization library enabling low-memory training alongside LoRA adapters.

Cost & effortWhat it takes

Low to moderate cost, often trainable on a single GPU in hours; needs a modest labeled dataset and comparatively little engineering effort versus full fine-tuning.

A living map of modern AI — kept current every morning