What Is RLHF (Reinforcement Learning from Human Feedback)?
RLHF is a training method that steers a model toward outputs humans prefer. People rank model responses; those rankings train a reward model that predicts human preference; and the language model is then optimised, by reinforcement learning, to score well under that reward model. It is how a raw next-token predictor becomes a helpful, instruction-following assistant — aligning the model with human preferences rather than with the raw statistics of its training text.
Preference, not a labelled right answer
Supervised fine-tuning needs a correct target for every input, which is impossible for open-ended generation — there is no single right essay. RLHF sidesteps that by learning from comparisons: humans need only say which of two responses is better, a far cheaper and more reliable judgement, and the reward model generalises those comparisons into a score the policy optimises against.
Its documented failure modes are the research frontier: the reward model is a proxy, and optimising hard against a proxy invites reward hacking and sycophancy — telling the user what scores well rather than what is true. Preference data also encodes the labellers’ biases. RLHF aligns a model to preferences; whether those preferences equal correctness is a separate, unsolved question.
Related standards
Questions
Is RLHF the same as fine-tuning?
It builds on fine-tuning but adds a learned reward and a reinforcement-learning stage; plain fine-tuning imitates fixed targets, RLHF optimises a preference score.
Does RLHF make a model truthful?
It makes it preferred, which is not the same — it can induce sycophancy, a known limitation driving work on alternatives.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.