Instruction Tuning
Supervised adaptation of a pretrained model on instruction-response examples so it follows task descriptions more reliably.
Machine-Learning Context
Instruction tuning adapts a pretrained model with supervised instruction-response examples so task descriptions and response formats are followed more reliably. Coverage across task families, data quality, instruction diversity, and consistent formatting strongly influence generalization.
Training Boundary
Instruction tuning is distinct from preference optimization. The former learns from supervised target responses, while approaches such as RLHF or DPO use comparative preference signals or reward-oriented objectives.
Related Machine-Learning Concepts
- Fine-Tuning
- RLHF
- DPO
- Supervised Fine-Tuning
What It Teaches—and What It Does Not
In instruction tuning, the supervised target is not merely a class label: a natural-language instruction, an input, and an expected response format can jointly define a training example. The FLAN work applied this idea across many task families and evaluated zero-shot generalization to unseen task types. The model is therefore trained not only on task content but also on the behavior of interpreting an instruction as a task specification.
This does not guarantee that every instruction-following answer is correct or safe. Instruction adherence, factual accuracy, and preference alignment remain distinct evaluation axes. Preference-based methods such as RLHF or DPO also use different training signals. For the direct experimental formulation of instruction tuning, see Wei et al., FLAN.