Train only a small set of continuous prompt vectors to steer a large frozen model, the lightest fine-tuning lever there is.
ConceptWhat it is
Prompt tuning keeps the entire base model frozen and learns only a small set of continuous embedding vectors, a soft prompt, that are prepended to the input. It sits at the lightest end of the PEFT family: where LoRA trains adapter matrices inside the layers, prompt tuning trains nothing but these few virtual tokens at the input embedding layer, making it one of the most parameter-minimal adaptation methods in common use.
Unlike hand-written prompt engineering, the soft prompt is not readable text; it is a block of learned vectors found by gradient descent, free to take values no real token maps to. Its headline property, established as model scale grew, is that it approaches full fine-tuning quality on very large models while staying weak on smaller ones, which is why it is often described as the cheapest lever available specifically on big models.
How it worksThe mechanics
A short sequence of trainable embedding vectors is initialized and prepended to the embeddings of every training example; the example flows through the frozen base model, a task loss is computed on the output, and gradients update only the soft-prompt vectors while every model weight stays untouched. After training, the learned soft prompt is saved as a tiny per-task file and, at inference, simply prepended again to steer the shared frozen model toward that task, so switching tasks means swapping one small soft prompt rather than loading a different model.
At a glanceSee it
Placing prompt tuning among its PEFT cousins by where each method injects new parameters — it sits at the lightest, input-embedding end while capacity and storage cost grow toward full fine-tuning.
A troubleshooting path for when a soft prompt underperforms — lengthen it, add layer-wise capacity, fix initialization, or accept that injecting new knowledge needs a different lever entirely.
When to use itWhere it fits
- Adapting a single very large frozen model to many narrow tasks, training and storing one tiny soft prompt for each.
- Small labeled datasets, where training a minimal parameter count limits the risk of overfitting.
- Multi-tenant or multi-task serving, where per-task soft prompts can be swapped cheaply on one shared base.
- When even LoRA feels heavy and the base model is large enough to respond well to input-only steering.
When NOT to use itLimits & anti-patterns
- Smaller base models, where prompt tuning underperforms and heavier methods like LoRA are worth their extra cost.
- Tasks needing new factual knowledge the base model lacks, which retrieval or fuller fine-tuning addresses better.
- Deep capability or behavior shifts that a handful of input vectors are simply not expressive enough to encode.
- Cases where a readable, auditable prompt matters, since a soft prompt is opaque learned vectors, not text.
Trade-offsAdvantages & costs
Advantages
- Trains the fewest parameters of any common adaptation method, so compute and storage cost stay very low.
- One frozen base serves unlimited tasks; each task is just a small soft-prompt file to swap in.
- Very low risk of catastrophic forgetting, since no model weight is ever changed.
- Fast, cheap iteration that fits on modest hardware for many practical model sizes.
Trade-offs & costs
- Weak on smaller models; its quality only rivals full fine-tuning as the base grows into the billions of parameters.
- Adds no new knowledge and only partially reshapes behavior, capping how far it can move the model.
- The soft prompt is opaque and not human-readable, making it harder to inspect or debug than a text prompt.
- Prepended virtual tokens consume context-window space and add a small per-request overhead.
ExampleIn the real world
Prompt tuning was introduced by Google researchers on the T5 model family, where they showed that as the base grew toward the multi-billion-parameter range, tuning only a soft prompt matched full fine-tuning on standard benchmark tasks. A team could reproduce the pattern today with Hugging Face PEFT: freeze one large open-weight model, train a separate soft prompt for each of several narrow classification tasks such as intent routing, topic labeling, and sentiment scoring, and serve them all from the single shared model by swapping in the relevant small soft prompt per request, never holding more than one copy of the weights in memory.
ToolsHow to implement it
- Hugging Face PEFTprovides PromptTuningConfig to attach and train a soft prompt on any supported base model, alongside prefix tuning and P-tuning.
- Google prompt-tuning repothe original T5X-based reference implementation released by the method's authors.
- OpenPromptopen-source research framework for defining and training continuous and discrete prompt-learning methods.
Cost & effortWhat it takes
The cheapest adaptation lever there is: the soft prompt holds only its token count times the model's embedding width, a tiny fraction of the base weights, so runs are fast and fit on modest hardware, and each new task adds just a small soft-prompt file rather than a full model copy. The real cost is conditional value, since it pays off mainly on large base models, plus the modest but genuine effort of assembling a clean task dataset and standing up the tuning loop.