Home › Fine-Tuning & Alignment › Instruction tuning
🎯 · Models

Instruction tuning

Teach a base model to follow instructions instead of just predicting next tokens.

In one line

Supervised fine-tuning on instruction-response pairs turns a text predictor into an assistant.

ConceptWhat it is

Instruction tuning is supervised fine-tuning on a dataset of instructions paired with desired responses, converting a raw next-token-prediction base model into one that reliably follows commands, answers questions, and adopts a helpful assistant behavior. It exists because base models trained only to predict the next token do not inherently know to behave like an assistant when given a prompt.

It is typically the first alignment step applied after pretraining, before further alignment stages like RLHF or DPO refine tone and preferences.

How it worksThe mechanics

A dataset of diverse instructions paired with high-quality responses, often written or curated by humans or generated and filtered by stronger models, is used to continue training the base model with standard supervised cross-entropy loss, teaching it the pattern of responding helpfully to a prompt rather than merely continuing it.

At a glanceSee it

Instruction tuning diagram
Instruction tuning diagram 1

Zoom into the training step — SFT joins prompt and response into one sequence, then masks the instruction so cross-entropy updates the model on the reply tokens alone.

Instruction tuning diagram 2

Where the pairs come from — humans, a stronger model, or recast public tasks — feeding a curation loop that prizes a small, diverse, high-quality set over raw volume.

When to use itWhere it fits

  • Converting a raw pretrained base model into a usable assistant for the first time.
  • Adapting an existing assistant to a specific response format or house style, like structured JSON outputs.
  • Teaching domain-specific instruction patterns, like following a customer support script.
  • As the first stage before applying RLHF or DPO for finer-grained preference alignment.

When NOT to use itLimits & anti-patterns

  • The model already follows instructions well and the actual problem is factual grounding, which RAG addresses better.
  • Fine-grained preference shaping, like tone or safety trade-offs, where DPO or RLHF are more direct tools than plain SFT.
  • Very small tweaks in behavior, where prompting or few-shot examples achieve the same result without any training.

Trade-offsAdvantages & costs

Advantages
  • Reliably transforms a base model into a usable, instruction-following system.
  • Well-understood, standard supervised training with mature tooling.
  • Can be combined with LoRA for a cheap, fast adaptation pass.
  • Establishes the foundation that later alignment stages build on.
Trade-offs & costs
  • Quality is capped by the quality and diversity of the instruction dataset.
  • Does not directly optimize for nuanced human preferences the way RLHF or DPO do.
  • Can overfit to the exact style of the training examples if the dataset is narrow.
  • Still requires curating or generating a solid labeled dataset, which is real engineering work.

ExampleIn the real world

The original Stanford Alpaca project instruction-tuned a Llama base model on 52,000 generated instruction-response pairs, turning it from a raw completion model into one that follows commands.

ToolsHow to implement it

  • Hugging Face TRLSFTTrainer utility purpose-built for instruction tuning workflows.
  • OpenAI fine-tuning APImanaged instruction tuning without owning training infrastructure.
  • Alpaca and Dolly datasetswidely used open instruction-response datasets for SFT experiments.
  • Axolotlconfig-driven framework for running instruction tuning at various model sizes.

Cost & effortWhat it takes

Moderate cost depending on model size and dataset volume; needs a curated instruction dataset of at least a few thousand examples, with engineering effort mostly in dataset quality control.

A living map of modern AI — kept current every morning