SFT trains a model on curated instruction-and-response pairs so it learns to be helpful.
DefinitionWhat it means
Supervised Fine-Tuning, SFT, continues training a pretrained base model on a labeled dataset of instruction-response pairs, using standard supervised loss so the model learns to produce the demonstrated response style for a given prompt. It is typically the first adaptation step, turning a raw next-token predictor into a model that follows instructions and holds a conversation.
Why it mattersWhy you should care
SFT is the foundation every instruction-tuned or chat model is built on; the quality, diversity, and correctness of its demonstration data sets a ceiling on everything that comes after, including RLHF and DPO. Teams building domain-specific assistants often run a lightweight SFT pass on top of a general instruct model before layering on preference optimization.
At a glanceSee it
The inner training loop the box diagram hides — SFT masks the prompt tokens so cross-entropy loss lands only on the response, teaching the model to answer rather than parrot the question back.
The data-curation decision behind the flat arrow — pairs whose answers sit outside the base model's knowledge train it to guess confidently, while an over-narrow set overfits style and erodes pretrained skills.
Where you see itIn the wild
- The step between a raw pretrained checkpoint and a usable chat model.
- Instruction datasets built from human-written or model-generated demonstrations.
- Domain adaptation projects fine-tuning an instruct model on internal support tickets.