RLHF
Reinforcement Learning from Human Feedback, a model-alignment approach that uses human preference signals to optimize model behavior through a reward-driven training stage.
Machine-Learning Context
RLHF uses human preference information to shape model behavior through a reward-oriented training process. A common pipeline includes supervised fine-tuning, preference collection or reward modeling, and policy optimization, with goals that may include helpfulness, format adherence, and safety rather than task accuracy alone.
Training Boundary
RLHF is not one fixed algorithm. Direct preference methods such as DPO use a different optimization formulation, while reward models and policy optimization introduce failure modes such as reward hacking and preference-model bias.
Related Machine-Learning Concepts
- DPO
- Instruction Tuning
- Preference Model
- Reward Model
Boundary of the Preference Signal
A common RLHF pipeline first establishes behavior with supervised fine-tuning, then collects human comparisons, may train a reward model to approximate those preferences, and optimizes a policy against that reward. The InstructGPT work is a well-documented example of this multi-stage design.
Human preference is not a mathematical oracle for factual truth. Annotator distribution, instructions, example selection, and reward-model error can all shape the learned objective; a higher reward is therefore not proof of real-world correctness. Methods such as DPO can also use preference data without requiring the same explicit reward-model-plus-RL stage, so not every preference-optimization method should be collapsed into RLHF. Source: Ouyang et al..