Model customization

Reinforcement Learning from Human Feedback

RLHF

A training technique in which people rate or compare a model's answers and the model is then adjusted to produce the kind of answers people preferred.

THE THREE STAGES OF RLHF1 · Comparepeople pick the better answer2 · Reward modellearns to predict the preference3 · Reinforcethe model is tuned to score high40,000 human choices become a reward model that can score millions of answers (illustrative figure).

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/rlhf

In plain terms

A pre-trained language model is like someone who has read an entire library and never held a conversation: knowledgeable, and just as likely to continue your question with another question as to answer it. RLHF is the coaching that follows. People look at pairs of answers and say which is better: more helpful, more accurate, more appropriate. The model is then nudged, over many rounds, towards the kind of answer that wins.

Why it matters

RLHF is the step that turned raw text predictors into assistants, and it explains much about how they behave: polite, cautious about some topics, inclined to give structured answers. Those traits were chosen, through the instructions given to the people who rated the answers. That is its limit as well. The model learns what raters approve of, which is close to what is good and differs from it in places: an answer that sounds confident and agreeable tends to be rated well. Flattery and overconfidence in assistants are partly side effects of this training. Buyers inherit these choices with the model.

Example

A lab has a pre-trained model that often rambles or ignores the question. It collects 40,000 prompts, generates two answers to each, and has trained raters choose the better one. A second model learns from those 40,000 choices to predict which answer people will prefer. The language model is then trained against that scorer over millions of answers. In blind tests afterwards, users prefer the new model's answer 85% of the time.

Most often confused with

RLHF vs. Fine-tuning

RLHFLearns from judgements: which of two answers is better
Fine-tuningLearns from demonstrations: here is the right answer

Ordinary fine-tuning is supervised: each example contains the ideal output and the model learns to reproduce it. RLHF needs no ideal output; raters only have to recognise the better answer, which is far easier than writing a perfect one, and the model can end up better than any example it was shown. RLHF is mostly done by model developers to shape general behaviour. Company projects mostly use supervised fine-tuning.

Origin: The method was described for reinforcement-learning agents by Paul Christiano and colleagues in 2017 and applied to large language models in OpenAI's InstructGPT paper of 2022; ChatGPT was trained with it.

Under the hood

The published recipe has three stages. First, supervised fine-tuning on human-written demonstrations. Second, a reward model is trained on human comparisons of model outputs, so that it can score any answer the way raters would. Third, the language model is optimised with a reinforcement-learning algorithm, usually PPO, to produce answers that the reward model scores highly, with a penalty that keeps it close to its starting point. Known weaknesses: rater disagreement and bias, reward models that can be gamed, sycophancy, and the cost of expert raters. Alternatives and successors: direct preference optimisation (DPO), which learns from the same comparisons without a separate reward model or a reinforcement-learning loop; feedback generated by a model following written principles, as in Anthropic's Constitutional AI; and reinforcement fine-tuning against automated graders. Current models are typically shaped by a combination of these, and labs do not publish every detail.

Written by Mehmet Erkek · Last updated: