Adapters add tiny bottleneck modules inside a frozen model so each task gets its own swappable, cheaply trained plug-in.
ConceptWhat it is
Adapters are small bottleneck layers inserted between the existing blocks of a pretrained transformer. They belong to the PEFT (parameter-efficient fine-tuning) family: the original model's weights stay frozen and only the newly added modules are trained. Each module projects the hidden state down to a small dimension, applies a nonlinearity, projects it back up, and adds the result back through a residual connection, so a freshly initialized adapter starts near identity and barely disturbs the base model.
Adapters exist because full fine-tuning is expensive and produces a full-size model copy for every task. Because the base is untouched, a trained adapter is a tiny swappable module — often just a few megabytes — that can be loaded, stacked, or composed at serving time. That makes adapters a natural fit for modular multi-task serving, where one shared base hosts many task-specific plug-ins instead of many full models.
How it worksThe mechanics
During training the base weights are frozen and a pair of small down- and up-projection layers is inserted after the attention and feed-forward sub-layers of each block; only these projections, and their layer norms, receive gradients. The bottleneck dimension sets the parameter budget — smaller means fewer parameters and less capacity to absorb the task. At inference the adapter runs inline as extra sequential layers, so its output is recomputed on every forward pass rather than folded back into the original weights, and switching tasks is just a matter of pointing the runtime at a different adapter file on top of the same frozen base.
At a glanceSee it
What the pipeline's single adapter node hides — down-project to a tiny rank r, a nonlinearity, up-project, then a residual add back to the input — and why only a sliver of parameters ever train.
When to choose classic bottleneck adapters over LoRA or full fine-tuning — the call turns on whether you can tolerate inline latency and must serve many tasks modularly from one loaded base.
When to use itWhere it fits
- Serving many task- or customer-specific variants from one shared base, where storing a full model per task is impractical.
- Building a library of reusable, swappable skills that can be added or removed without retraining the base.
- Moderate labeled data per task and a low compute budget, where full fine-tuning is overkill.
- Composing several capabilities at once, for example fusing separately trained domain and task adapters.
When NOT to use itLimits & anti-patterns
- Latency-critical serving, since adapters add sequential compute on every forward pass and cannot be merged away the way LoRA can.
- Deep changes to the model's core knowledge or reasoning, which small bottleneck modules cannot deliver.
- Single-task deployments, where a mergeable method like LoRA reaches similar quality with zero added latency.
- When the base simply lacks the needed facts — retrieval or continued pretraining fits better than any adapter.
Trade-offsAdvantages & costs
Advantages
- Trains only a tiny fraction of parameters, so runs finish fast on modest hardware.
- Each adapter is a few megabytes, cheap to store, version, and ship independently of the base.
- Strong modularity — swap, stack, or fuse adapters at serving time without touching the shared base.
- The frozen base means little catastrophic forgetting and a clean rollback by simply removing the module.
Trade-offs & costs
- Adds real inference latency because the extra layers run inline and cannot be folded into the base weights.
- The bottleneck size is a capacity knob that needs tuning; set too small, the adapter underfits the task.
- Lower performance ceiling than full fine-tuning for large distribution shifts.
- Still needs a labeled dataset per task, plus routing plumbing to attach the right adapter to each request.
ExampleIn the real world
A support platform serves dozens of enterprise customers from one open-weight base model. Rather than fine-tuning and hosting a separate full-size model per customer, the team trains one bottleneck adapter per customer on that customer's ticket history and tone guidelines, each only a few megabytes. At request time the router loads the base once and attaches the caller's adapter, so a single GPU fleet serves every customer; onboarding a new customer means training and dropping in one more adapter file, and the small extra latency from the inline modules is an acceptable trade for the chat use case.
ToolsHow to implement it
- AdapterHuband its
adapterslibrary (formerly adapter-transformers) — the canonical bottleneck-adapter implementation plus a hub of pretrained adapters. - Hugging Face PEFTa parameter-efficient fine-tuning library implementing LoRA, IA3, and related adapter-style methods on top of frozen models.
- AdapterFusiontechnique from the AdapterHub group for combining several trained adapters for multi-task inference.
- LLaMA-Adaptera lightweight adapter approach for efficiently adapting LLaMA-family models.
Cost & effortWhat it takes
Low compute: because only the inserted modules train, jobs typically run on a single GPU in hours and need only a moderate labeled dataset per task. Storage is trivial — adapters are measured in megabytes, not gigabytes — which is exactly what makes many-task serving affordable. The recurring cost sits elsewhere: the inline modules levy a small but permanent latency and compute tax on every request, and operating a multi-adapter fleet adds routing and lifecycle plumbing to build and maintain.