RLHF

Turkish equivalent: İnsan geri bildirimiyle pekiştirmeli öğrenmeDomain: Machine Learning

Reinforcement Learning from Human Feedback, a model-alignment approach that uses human preference signals to optimize model behavior through a reward-driven training stage.

Machine-Learning Context

RLHF uses human preference information to shape model behavior through a reward-oriented training process. A common pipeline includes supervised fine-tuning, preference collection or reward modeling, and policy optimization, with goals that may include helpfulness, format adherence, and safety rather than task accuracy alone.

Training Boundary

RLHF is not one fixed algorithm. Direct preference methods such as DPO use a different optimization formulation, while reward models and policy optimization introduce failure modes such as reward hacking and preference-model bias.

Boundary of the Preference Signal

A common RLHF pipeline first establishes behavior with supervised fine-tuning, then collects human comparisons, may train a reward model to approximate those preferences, and optimizes a policy against that reward. The InstructGPT work is a well-documented example of this multi-stage design.

Human preference is not a mathematical oracle for factual truth. Annotator distribution, instructions, example selection, and reward-model error can all shape the learned objective; a higher reward is therefore not proof of real-world correctness. Methods such as DPO can also use preference data without requiring the same explicit reward-model-plus-RL stage, so not every preference-optimization method should be collapsed into RLHF. Source: Ouyang et al..