Supervised fine-tuning on instruction-response pairs turns a text predictor into an assistant.
ConceptWhat it is
Instruction tuning is supervised fine-tuning on a dataset of instructions paired with desired responses, converting a raw next-token-prediction base model into one that reliably follows commands, answers questions, and adopts a helpful assistant behavior. It exists because base models trained only to predict the next token do not inherently know to behave like an assistant when given a prompt.
It is typically the first alignment step applied after pretraining, before further alignment stages like RLHF or DPO refine tone and preferences.
How it worksThe mechanics
A dataset of diverse instructions paired with high-quality responses, often written or curated by humans or generated and filtered by stronger models, is used to continue training the base model with standard supervised cross-entropy loss, teaching it the pattern of responding helpfully to a prompt rather than merely continuing it.
At a glanceSee it
Zoom into the training step — SFT joins prompt and response into one sequence, then masks the instruction so cross-entropy updates the model on the reply tokens alone.
Where the pairs come from — humans, a stronger model, or recast public tasks — feeding a curation loop that prizes a small, diverse, high-quality set over raw volume.
When to use itWhere it fits
- Converting a raw pretrained base model into a usable assistant for the first time.
- Adapting an existing assistant to a specific response format or house style, like structured JSON outputs.
- Teaching domain-specific instruction patterns, like following a customer support script.
- As the first stage before applying RLHF or DPO for finer-grained preference alignment.
When NOT to use itLimits & anti-patterns
- The model already follows instructions well and the actual problem is factual grounding, which RAG addresses better.
- Fine-grained preference shaping, like tone or safety trade-offs, where DPO or RLHF are more direct tools than plain SFT.
- Very small tweaks in behavior, where prompting or few-shot examples achieve the same result without any training.
Trade-offsAdvantages & costs
Advantages
- Reliably transforms a base model into a usable, instruction-following system.
- Well-understood, standard supervised training with mature tooling.
- Can be combined with LoRA for a cheap, fast adaptation pass.
- Establishes the foundation that later alignment stages build on.
Trade-offs & costs
- Quality is capped by the quality and diversity of the instruction dataset.
- Does not directly optimize for nuanced human preferences the way RLHF or DPO do.
- Can overfit to the exact style of the training examples if the dataset is narrow.
- Still requires curating or generating a solid labeled dataset, which is real engineering work.
ExampleIn the real world
The original Stanford Alpaca project instruction-tuned a Llama base model on 52,000 generated instruction-response pairs, turning it from a raw completion model into one that follows commands.ToolsHow to implement it
- Hugging Face TRLSFTTrainer utility purpose-built for instruction tuning workflows.
- OpenAI fine-tuning APImanaged instruction tuning without owning training infrastructure.
- Alpaca and Dolly datasetswidely used open instruction-response datasets for SFT experiments.
- Axolotlconfig-driven framework for running instruction tuning at various model sizes.
Cost & effortWhat it takes
Moderate cost depending on model size and dataset volume; needs a curated instruction dataset of at least a few thousand examples, with engineering effort mostly in dataset quality control.