DPO
DPO — Direct Preference Optimization, a training objective that learns from preference pairs without explicitly fitting a separate reward model and reinforcement-learning loop.
Machine-Learning Context
Direct Preference Optimization learns from chosen/rejected response pairs using an objective related to a reference policy, avoiding the separate reward-model and reinforcement-learning loop used by some RLHF pipelines. It can simplify preference optimization while preserving a direct connection to preference data.
Generalization Boundary
DPO is not universally superior for every alignment problem. Preference quality, reference-policy behavior, the beta parameter, data coverage, and evaluation method all affect the resulting policy.
Related Machine-Learning Concepts
- RLHF
- Preference Learning
- Instruction Tuning
- Fine-Tuning
Source
- https://arxiv.org/abs/2305.18290