DPO
Direct Preference Optimization, a training objective that learns from preference pairs without explicitly fitting a separate reward model and reinforcement-learning loop.
Related Concepts
- RLHF
- Preference Learning
- Instruction Tuning
- Fine-Tuning
Source
- https://arxiv.org/abs/2305.18290