Home › Fine-Tuning & Alignment › Prefix tuning
🎯 · Models

Prefix tuning

Prefix tuning: train small key and value vectors at every attention layer to steer behaviour cheaply.

In one line

Prefix tuning freezes the base model and trains only tiny prefix vectors injected into each layer's attention, changing how the model behaves without changing what it knows.

ConceptWhat it is

Prefix tuning is a parameter-efficient fine-tuning (PEFT) method that adapts a large language model without touching any of its original weights. Instead of updating the network, it learns a short sequence of continuous prefix vectors that are prepended to the keys and values inside the self-attention of every transformer layer. Because these vectors sit in the attention computation at all depths, they can bias the model's behaviour throughout the forward pass while the frozen backbone does the heavy lifting.

It exists to make behavioural steering cheap and modular. A single hosted base model can serve many tasks or styles, each represented by its own small prefix that is swapped in per request. This is the deeper cousin of prompt tuning, which only edits soft tokens at the input embedding layer; by acting at every layer, prefix tuning is more expressive than prompt tuning while remaining far lighter than full fine-tuning. The trade-off is that it steers behaviour and format well but does not reliably inject new factual knowledge.

How it worksThe mechanics

Pick a prefix length (a handful to a few dozen virtual tokens) and initialize a trainable table of key and value vectors for each attention layer, commonly produced through a small reparameterization network during training for stability. Freeze all base-model weights. During each forward pass, these prefix key and value vectors are concatenated in front of the real tokens' keys and values, so every query attends to the prefix as if it were extra context. Run the task examples, compute the loss on the model's outputs, and backpropagate gradients only into the prefix parameters, leaving the backbone untouched. After training, discard the reparameterization network, keep only the resulting prefix vectors as a tiny artifact, and at inference load the prefix matching the target task or style; switching behaviour is just switching which prefix is attached.

At a glanceSee it

Prefix tuning diagram
Prefix tuning diagram 1

Direct prefix optimization is unstable, so training routes through a reparameterization MLP that is thrown away at deploy time — only the learned key and value vectors ship.

Prefix tuning diagram 2

Situating prefix tuning among its PEFT siblings by where each injects trainable parameters — its edge is reaching every layer's attention, not just the input.

When to use itWhere it fits

  • You want to steer tone, persona, format, or task behaviour of a frozen model rather than teach it new facts.
  • One base model must serve many tasks and you need cheap, hot-swappable adapters with near-zero storage cost per task.
  • Compute and labeled data are limited, so full fine-tuning is impractical but soft prompting alone is too weak.
  • You need to keep the original model byte-for-byte intact for governance, rollback, or shared-tenant reasons.

When NOT to use itLimits & anti-patterns

  • The goal is to add new domain knowledge or facts the base model never learned; pair retrieval or heavier tuning instead.
  • You need maximum task accuracy or a large behaviour shift, where LoRA or full fine-tuning tends to be more expressive.
  • The task demands very long user inputs and you cannot spare context budget, since the prefix consumes attention positions.
  • You want a single merged model with no inference-time adapter logic; prefix vectors must be attached at runtime.

Trade-offsAdvantages & costs

Advantages
  • Trains only a tiny fraction of parameters, so it is fast and memory-light to fit and to store.
  • Fully modular: many prefixes coexist against one frozen backbone and swap in per request.
  • The base model is never modified, preserving its behaviour and simplifying rollback and multi-tenant serving.
  • Acts at every layer, giving stronger, more stable steering than input-only prompt tuning.
Trade-offs & costs
  • Less expressive than LoRA or full fine-tuning, so hard tasks may hit an accuracy ceiling.
  • Poor at injecting new factual knowledge; it shapes behaviour, not what the model knows.
  • The prefix occupies attention positions, trimming usable context and adding a little inference overhead.
  • Training can be sensitive to prefix length and initialization, often needing a reparameterization trick to converge well.

ExampleIn the real world

A support team hosts one frozen mid-size instruction model and needs three consistent response styles: a formal, compliance-careful voice for regulated products, a concise casual voice for consumer apps, and a brief bilingual voice for a second market. Rather than fine-tune three separate models, they train one prefix per style on a few thousand curated dialogues each, backpropagating only into the prefix vectors. Each learned prefix is a small artifact, so all three fit alongside the single shared model. At serving time the router picks the prefix matching the product line and attaches it before generation, so tone and formatting shift on demand. Because the prefixes only steer behaviour, the model still cannot cite policies it never saw, so the team keeps a retrieval layer for the actual policy text and lets the prefix govern how that content is phrased.

ToolsHow to implement it

  • Hugging Face PEFT, which implements prefix tuning alongside LoRA and prompt tuning behind one adapter API.
  • P-Tuning v2, the deep-prefix variant designed to make this approach competitive on understanding tasks across model sizes.
  • OpenDelta, a library for attaching and managing delta or adapter modules on frozen transformers.
  • PyTorch as the underlying training framework for defining and optimizing the prefix parameters.

Cost & effortWhat it takes

Cost and effort are low. Training touches only the prefix, so a single modest GPU and a moderate labeled set are usually enough, and runs finish quickly relative to full fine-tuning. Storage per task is negligible, well under one percent of the base parameters, which makes maintaining many tasks cheap. The recurring costs are engineering discipline around prefix length and initialization tuning, a small inference-time overhead from the added attention positions, and the runtime plumbing to load and swap the correct prefix per request.

A living map of modern AI — kept current every morning