What Is RLHF (Reinforcement Learning from Human Feedback)?

RLHF is a training method that steers a model toward outputs humans prefer. People rank model responses; those rankings train a reward model that predicts human preference; and the language model is then optimised, by reinforcement learning, to score well under that reward model. It is how a raw next-token predictor becomes a helpful, instruction-following assistant — aligning the model with human preferences rather than with the raw statistics of its training text.

Signal
Human rankings of model outputs
Steps
Collect preferences → reward model → RL optimisation
Turns
A text predictor into an instruction-follower

Preference, not a labelled right answer

Supervised fine-tuning needs a correct target for every input, which is impossible for open-ended generation — there is no single right essay. RLHF sidesteps that by learning from comparisons: humans need only say which of two responses is better, a far cheaper and more reliable judgement, and the reward model generalises those comparisons into a score the policy optimises against.

Its documented failure modes are the research frontier: the reward model is a proxy, and optimising hard against a proxy invites reward hacking and sycophancy — telling the user what scores well rather than what is true. Preference data also encodes the labellers’ biases. RLHF aligns a model to preferences; whether those preferences equal correctness is a separate, unsolved question.

Related standards

Christiano et al., 2017 — Deep RL from Human Preferences

Questions

Is RLHF the same as fine-tuning?

It builds on fine-tuning but adds a learned reward and a reinforcement-learning stage; plain fine-tuning imitates fixed targets, RLHF optimises a preference score.

Does RLHF make a model truthful?

It makes it preferred, which is not the same — it can induce sycophancy, a known limitation driving work on alternatives.

Keep reading

related
What Is Fine-Tuning?
related
What Is Chain-of-Thought Prompting?
related
What Are AI Guardrails?
related
What Is Model Evaluation (Benchmarks and Evals)?
referenced by
Supervised vs Unsupervised Learning: What’s the Difference?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
Managing an Agent’s Context Window
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is a Mixture-of-Experts Model?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an Embedding?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is a Transformer (Neural Network Architecture)?
Agent Data & Memory
What Is the Attention Mechanism?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Model Distillation?
Agent Data & Memory
What Is Model Quantization?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
Transformer vs RNN: What Changed?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.