Preference Optimization: RLHF, DPO & GRPO
How a model goes from producing plausible text to producing text people want. Four methods, one objective, and a steadily shrinking amount of machinery needed to reach it.
Supervised fine-tuning teaches format but cannot teach preference, because preference is a comparison between two acceptable answers and a label on one answer cannot express it. Everything on this page exists to learn from comparisons.
All four methods optimise the same objective: raise reward while staying close to the starting model, measured by KL divergence. They differ only in how much machinery they need. RLHF trains a reward model and runs PPO against it, holding four models in memory.
DPO proves the reward model cancels algebraically and trains on preference pairs directly. GRPO keeps the RL loop but replaces the value model with the mean reward of a sampled group. RLVR throws out the learned reward model entirely and scores against a checker that is right.
One sentence for an interview: the field spent three years discovering how much of the RLHF pipeline was removable, and the answer turned out to be most of it.
01Why supervised fine-tuning is not enough
The limitation comes from the structure of the objective, and more scale or better data quality does not remove it.
Page 10 covers SFT: show the model prompt-and-response pairs, train on next-token prediction, and it learns to answer in the right shape. That works because there is a target to imitate.
Now consider a question with two good answers. One is concise, one is thorough. One hedges appropriately, one commits. Both are fluent, both are correct, and a human asked to choose would have a clear opinion. SFT has no way to express that opinion, because its loss function only knows how to raise the likelihood of one specific token sequence.
SFT can only ever tell a model "produce this". It cannot tell it "produce this and not that". Preference is a relation between two outputs, and you cannot encode a relation in a per-token label on one of them.
There is a second problem, and it is the more practical one. Training on demonstrations pushes the model toward the average of its demonstrators. If the people writing your SFT data are good but not exceptional, imitation caps you at good but not exceptional.
Comparisons do not have that ceiling: people can reliably judge which of two answers is better long past the point where they could write the better one themselves.
That gap between judging quality and producing quality is what preference optimization converts into training signal.
02Intuition
Every method here is the same two-part bargain, with different amounts of machinery bolted on.
The bargain is:
- Move toward what people prefer. Some notion of reward goes up.
- Do not move far. Stay close to the model you started from.
The second half is not a technicality. It is the part that keeps the method from destroying the model. A policy optimising an imperfect reward with no constraint will find the exploits: repeating flattering phrases, padding length, adopting whatever surface pattern the reward signal happens to correlate with. It will score brilliantly and be useless. The distance constraint is what keeps the optimiser inside the region where the reward still means something.
The whole history of this field is a sequence of answers to one question: how much apparatus do you need to run that bargain?
Instruction-tuned models were trained purely by imitation. Quality was bounded by the demonstrations, and there was no mechanism to express that one acceptable answer was better than another.
Ouyang et al. (2022) trained a reward model on human comparisons and optimised against it with PPO. In the InstructGPT results, a 1.3B model tuned this way was preferred to a 175B model that was not.
Preference tuning became the default final stage of every serious model. The subsequent work has been steadily removing pieces of the pipeline while keeping the result.
03The pipeline, and where each method cuts into it
Read this as a single diagram in prose. Each method below deletes a different box.
- Pretraining. Next-token prediction on a large corpus. Produces a base model that can continue text but not follow instructions.
- SFT. It fine-tunes on demonstrations and produces a model that answers in the right shape. This checkpoint matters later, because it becomes the reference model that everything else is measured against.
- Reward modelling. Collect comparisons, train a model to score responses.
- Policy optimization. Update the SFT model to score well under that reward, penalised for drifting from the reference.
Now the cuts:
- RLHF runs all four stages as written.
- DPO deletes stage 3 and replaces stage 4 with a supervised loss on the preference pairs. The reward model does not just get skipped, it provably cancels.
- GRPO keeps stages 3 and 4 but deletes the value network inside stage 4, using the average reward of several sampled answers as the baseline instead.
- RLVR deletes stage 3 by replacing the learned reward model with a program that checks whether the answer is correct.
04The reward model, and the assumption underneath it
A reward model turns "A is better than B" into a number. The step from comparisons to a scalar is where an assumption enters, and it is worth naming.
You cannot ask annotators to score a response out of ten. People are inconsistent at absolute judgements and drift over a session. They are far more reliable at comparisons. So the data is pairs: a prompt $x$, a preferred response $y_w$, a dispreferred one $y_l$.
To get a scalar reward out of that, you need a model of how preferences relate to scores. The standard choice is Bradley-Terry, which says the probability of preferring one item is the sigmoid of the difference in their underlying scores:
Only the difference in rewards is identified. Adding a constant to every reward leaves every preference probability unchanged, which is why reward-model outputs have no absolute meaning and cannot be compared across training runs.
Fitting that by maximum likelihood over the dataset gives the reward model's training loss:
Architecturally the reward model is usually the SFT model with the token-prediction head swapped for a single scalar output, read off the final token.
Bradley-Terry assumes a single consistent scalar of quality exists, and that everyone is noisily measuring the same one. This is false in an interesting way: annotators disagree about tone, caution and verbosity, and averaging their disagreement produces a reward model that represents nobody's actual preference.
It also cannot represent intransitive preferences, where A beats B, B beats C, and C beats A. Both limitations are inherited by every method on this page that fits a Bradley-Terry reward, DPO included.
05RLHF with PPO: four models in the loop
The original recipe, and still the one the others are defined against.
With a reward model in hand, the objective is the bargain from section 02, written down:
Maximise reward, minus $\beta$ times the KL divergence from the reference model. $\beta$ is the exchange rate between the two, and it is the single most consequential hyperparameter in the whole pipeline.
Optimising this with PPO means the RL machinery from page 15 applies directly: the language model is the policy, generating a token is an action, and the sequence so far is the state. PPO's clipped surrogate objective prevents any single update from moving the policy too far, which page 15 derives.
In practice the KL term is not applied as a separate penalty on the objective. It is folded into the reward at each token, so the per-token reward becomes the reward model's score (at the end of the sequence) minus $\beta$ times the log-ratio between policy and reference at that token. The RL algorithm then optimises a single shaped reward.
The cost of all this is the thing people complain about. You hold four models at once:
- Policy. This is the model being trained, and it needs gradients and optimiser state.
- Reference. A frozen copy of the SFT checkpoint, for the KL term.
- Reward model. It is frozen and scores completions.
- Value model. It estimates expected return from a state, which is used to compute the advantage, and it is trained alongside the policy.
Two of those need gradients and two are frozen, but all four occupy memory, and the loop also requires generating rollouts, which is slow. The next three sections exist because of this cost.
06DPO: making the reward model cancel
The result that changed practice: you do not need the reward model, and this is not an approximation.
Page 10 states the DPO loss and explains each term. What follows is why it is true, because the derivation is short and it is the difference between using DPO and understanding when it will fail.
Derivation How the reward model disappears from the RLHF objective
The claim is that optimising the RLHF objective in section 05 is equivalent to a supervised loss on preference pairs. The path runs through the closed-form solution of the KL-constrained objective.
-
$$\max_{\pi}\ \mathbb{E}_{y \sim \pi}\big[r(x,y)\big] - \beta\, D_{KL}\big(\pi \,\|\, \pi_{\text{ref}}\big)$$Where we start: the RLHF objective, over all possible policies rather than a parameterised family. This is the step that makes the result exact but also marks its limit, since a real network cannot represent every distribution.
-
$$\pi^\star(y \mid x) = \frac{1}{Z(x)}\, \pi_{\text{ref}}(y \mid x)\, \exp\!\Big(\tfrac{1}{\beta} r(x,y)\Big)$$Why: this objective has a known closed-form maximiser. It is the reference distribution reweighted by exponentiated reward, a Gibbs distribution, with $Z(x) = \sum_y \pi_{\text{ref}}(y \mid x)\exp(r(x,y)/\beta)$ normalising it. Computing $Z(x)$ requires summing over every possible response, so this formula cannot be used directly. That is the whole obstacle.
-
$$r(x,y) = \beta \log \frac{\pi^\star(y \mid x)}{\pi_{\text{ref}}(y \mid x)} + \beta \log Z(x)$$Why: rearrange the previous line for $r$. Read it in this direction and it says something surprising: any policy implicitly defines a reward function, namely its log-ratio against the reference. The intractable $Z(x)$ is still here, but note it depends only on the prompt.
-
$$P(y_w \succ y_l \mid x) = \sigma\big(r(x,y_w) - r(x,y_l)\big)$$Why: bring in Bradley-Terry from section 04, the same model the reward model was fitted with. Preferences depend only on a difference of rewards, which is what the next step exploits.
-
$$r(x,y_w) - r(x,y_l) = \beta \log \frac{\pi^\star(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi^\star(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}$$The step that does the work: substitute step 3 into step 4. Both responses share the same prompt $x$, so both carry the identical $\beta \log Z(x)$ term, and subtracting removes it. The intractable normaliser is gone, and with it the need for a reward model at all.
-
$$\mathcal{L}_{\text{DPO}} = -\,\mathbb{E}_{(x,y_w,y_l)}\left[\log \sigma\!\left(\beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}\right)\right]$$Where we land: maximum likelihood on the preference data under this reparameterisation, with $\pi_\theta$ in place of the unknown optimum. Every quantity is a log-probability the policy can compute directly. No reward model, no rollouts, no value network.
One consequence worth internalising: the quantity $\beta \log \frac{\pi_\theta(y \mid x)}{\pi_{\text{ref}}(y \mid x)}$ is DPO's implicit reward. You never train a reward model, but you can still read one off the policy at any point, and the gap between the implicit rewards of the chosen and rejected responses is the margin. Watching that margin is how you tell whether training is working, and section 09 is built around it.
07GRPO: dropping the value model too
DPO removes the RL loop. GRPO keeps it and removes something else, which turns out to matter enormously for reasoning.
Group Relative Policy Optimization was introduced in the DeepSeekMath paper and is the method behind DeepSeek-R1. It starts from a question about PPO: what is the value model for?
Its job is to provide a baseline. Advantage is reward minus expected reward, and subtracting a baseline is what stops every action in a good episode from being reinforced equally. PPO learns that baseline with a second network roughly the size of the policy.
GRPO's observation is that for language you can get a baseline for free. Sample a group of $G$ responses to the same prompt, score them all, and use the group's own statistics:
The advantage of response $i$ is just how much better it was than its siblings, in units of the group's own spread. No learned value function anywhere.
That advantage then goes into a PPO-style clipped objective with a KL term to the reference. The clipping and the KL constraint are unchanged; only the source of the baseline differs.
For a maths or code problem you can sample sixteen attempts and grade every one automatically. The group mean is then a good estimate of how hard this specific prompt is, which is exactly what a value model spends its capacity trying to learn.
Normalising within the group also makes easy and hard prompts contribute comparably, and easy prompts do not dominate because of higher absolute rewards. You get a per-prompt difficulty baseline for the price of sampling, and you delete a network the size of your policy.
The cost is generation: every optimisation step needs $G$ full completions per prompt and not one, so GRPO is inference-heavy in a way DPO is not. Practically this means the throughput ideas on page 31 apply directly to your training loop, and serious GRPO setups run a dedicated inference engine to produce rollouts.
The DeepSeek-R1 work pushed this further than most expected. Its central claim is that reasoning ability can be incentivised through reinforcement learning alone, without human-labelled reasoning traces to imitate. The model is rewarded for reaching correct answers and develops the intermediate reasoning by itself.
08RLVR: when the reward is a checker, not a model
The last piece to remove is the learned reward model, and for some tasks you can delete it outright.
A learned reward model is a stand-in for a human judgement you cannot query at training speed. But for a whole class of tasks there is no judgement required, only a fact:
- A maths answer either matches the reference or it does not.
- Code either passes the test suite or it does not.
- Output either parses as valid JSON against a schema or it does not.
Reinforcement learning with verifiable rewards uses that check as the reward directly. The Tulu 3 work applies it as a named stage in an open post-training recipe, and it is central to how recent reasoning models are trained.
What this buys is a reward you cannot hack in the usual way. A learned reward model is an imperfect proxy, and enough optimisation pressure finds the gap between the proxy and the thing it stands for. A test suite has no gap of that kind: passing the tests is the objective.
What it does not buy is safety from all gaming. A model can still special-case the visible tests, hard-code expected outputs, or find a degenerate answer format that the checker accepts. The pressure moves from exploiting a reward model's blind spots to exploiting a verifier's incompleteness, which is a narrower target but not an empty one.
The real limit is scope. Verifiable rewards work where correctness is decidable, which excludes most of what preference tuning was invented for. Nothing checks whether an answer was appropriately tactful. In practice the two are combined: verifiable rewards for reasoning and code, preference data for tone and helpfulness.
09Build this
DPO has a pathology you will not believe until you plot it, and plotting it takes one afternoon.
Run DPO on a small instruction-tuned model with a public preference dataset. Then log three quantities every step: the log-probability of the chosen response, the log-probability of the rejected one, and the margin between their implicit rewards. Most tutorials plot only the margin, and the margin is the least interesting of the three.
- Take a small instruction-tuned model, 0.5B to 1.5B, and a standard preference dataset with prompt, chosen and rejected fields. A free Colab GPU is enough at this size with LoRA.
- Set up
DPOTrainerfrom TRL with $\beta = 0.1$. With LoRA you can skip loading a separate reference model, since disabling the adapter recovers the base policy. - Log
logps/chosen,logps/rejectedand the reward margin every step, and plot all three on shared axes. - Train well past the point where the margin looks healthy, then read the two log-probability curves and not the margin.
- Now sweep $\beta$ with a much smaller value and a much larger one, and at each setting generate a few completions and read them.
10The variants
DPO spawned a large family. Most differ in one specific way, and knowing which is enough.
Identifies that DPO's sigmoid loss can be driven arbitrarily far when a pair is perfectly separable, which overfits the preference data. Replaces it with a bounded objective that cannot run away.
Relevant when preference pairs are clean and the model overfits them quickly.
Drops the need for pairs entirely. Learns from individual responses labelled good or bad, using an objective drawn from prospect theory.
Reach for it when you have thumbs-up/thumbs-down production feedback, which is far cheaper to collect than matched pairs.
Removes the reference model, and folds preference optimization into SFT as a single stage with an odds-ratio penalty on the rejected response.
Reach for it to halve memory and collapse two training stages into one.
Also reference-free, using length-normalised average log-probability as the implicit reward. The normalisation directly targets DPO's tendency to reward longer responses.
Reach for it when outputs are drifting longer without getting better.
A caution applies to all of them, including the numbers in their papers: preference-tuning results are usually reported on judge-scored benchmarks, and LLM judges have documented biases, length being the best known. A method that produces longer answers can post better numbers without producing better answers. Page 25 covers why that happens. Treat published win rates as a reason to try something on your own data and not as a ranking.
11What breaks
The failure modes are shared across methods, because they come from the objective and not from the algorithm.
- Reward hacking. The policy finds where the reward model is wrong and lives there. Classic symptom: reward climbing steadily while sampled outputs get worse. Always read generations; never trust the reward curve alone.
- Length bias. Human annotators mildly prefer longer answers, reward models learn that preference and amplify it, and the policy exploits it. Length is the single most common confound in this entire area.
- Both log-probabilities falling in DPO. This is the section 09 phenomenon, where the margin looks fine while the model becomes less likely to produce either response.
- KL collapse and lost diversity. Too much optimisation pressure and the model converges on one safe phrasing for everything. Benchmarks may not notice; users do.
- Reward model distribution shift. The reward model was trained on SFT-era outputs. As the policy improves it generates text unlike anything the reward model scored, and the scores become unreliable exactly when the policy is doing best.
- Preference data quality. Annotator agreement on subjective comparisons is often modest, and a reward model fitted to noisy labels is a noisy reward model. This is usually the binding constraint, more than the choice of algorithm.
Keep a fixed set of prompts and generate from them at every checkpoint, then read the outputs yourself. Every failure above is visible in generations long before it shows up in a metric, and several of them never show up in a metric at all.
12Where you meet this in the wild
Which method fits depends almost entirely on what kind of signal you can get.
You have a house style and some examples of good and bad replies. Start with SFT, then DPO on a few thousand pairs. The reward model is not worth building at this scale, and DPO is stable enough to run without an RL specialist.
Answers are checkable, so use GRPO with verifiable rewards. There is no reward model and no value model, and the signal is exactly right and not a proxy. Budget for generation: this is where most of the compute goes.
You have thumbs-up and thumbs-down on individual responses and no matched pairs. KTO is built for this shape of data, and constructing artificial pairs from unpaired signals usually loses more than it gains.
Full RLHF still appears here, because online methods can learn from the model's own current samples and not a fixed dataset. Constitutional AI adds a further step, using AI feedback against a written set of principles to reduce how much human labelling is needed.
13Interview questions
BeginnerWhy is supervised fine-tuning not enough to align a model?
Because SFT can only raise the likelihood of one target sequence, and preference is a relation between two acceptable outputs rather than a property of one. Given a concise answer and a thorough one that are both correct, SFT has no way to express that people prefer one. There is also a ceiling effect: imitation pushes the model toward the average of its demonstrators, whereas people can reliably judge which of two answers is better long past the point where they could write the better one. Preference optimization converts that judging ability into training signal.
BeginnerWhat is the KL term in the RLHF objective for?
It keeps the policy close to the SFT reference model, and it is what stops the method from destroying the model. The reward model is an imperfect proxy, so an unconstrained optimiser will find where the proxy is wrong and exploit it, producing text that scores highly and reads badly. The coefficient $\beta$ sets the exchange rate between reward and drift. Too low and the model degenerates; too high and it barely moves from where it started.
IntermediateHow does DPO eliminate the reward model, and is it an approximation?
It is not an approximation on the stated objective. The KL-constrained RLHF objective has a closed-form optimal policy, the reference distribution reweighted by exponentiated reward. Rearranging gives the reward as a log-ratio between policy and reference, plus a normalising term that depends only on the prompt. Substituting that into the Bradley-Terry preference model leaves a difference of two rewards for the same prompt, so the intractable normaliser cancels, and what remains is a supervised loss on the policy's own log-probabilities. The caveats are that the equivalence assumes an unconstrained policy class, and that DPO learns only from the fixed dataset rather than from its own samples.
IntermediateWhat does GRPO remove relative to PPO, and what does it cost?
It removes the value network. That network exists to provide a baseline for the advantage, and GRPO gets a baseline instead by sampling a group of responses to the same prompt and normalising each reward by the group's mean and standard deviation. That deletes a model roughly the size of the policy and gives a per-prompt difficulty baseline, which is exactly what the value function was trying to learn. The cost is generation: every step needs several full completions per prompt rather than one, so GRPO is inference-heavy and serious setups run a dedicated inference engine for rollouts.
IntermediateReward is climbing steadily during RLHF but the outputs look worse. What is happening?
Reward hacking. The policy has found a region where the reward model is wrong and is optimising the proxy rather than the thing it stands for. A related cause is distribution shift: the reward model was trained on SFT-era outputs, and as the policy moves away its scores become unreliable. Check whether the KL from the reference is growing, look for the usual exploits such as length inflation or repeated flattering phrasing, and read generations from a fixed prompt set at every checkpoint. Remedies are raising $\beta$, stopping earlier, or refreshing the reward model on samples from the current policy.
DeepIn DPO, both chosen and rejected log-probabilities decrease during training. Is that a bug?
It is a real and commonly observed property of the objective rather than an implementation error. The loss depends only on the difference between the two implicit rewards, so it is satisfied equally well by raising the chosen response's likelihood or by lowering the rejected one further, and gradient descent frequently does the latter to both. The consequence is that the model can become less likely to produce the preferred response while the margin looks healthy. It is a good argument for monitoring the raw log-probabilities rather than only the margin, and part of the motivation for reference-free variants and for online methods that sample from the current policy.
DeepWhat does the Bradley-Terry assumption commit you to, and when does it fail?
It assumes a single scalar of quality exists and that every annotator noisily measures the same one, with preference probability given by the sigmoid of a reward difference. Two consequences follow. Only differences are identified, so reward-model outputs have no absolute meaning and cannot be compared across runs. More seriously, genuine disagreement about tone, caution or verbosity gets averaged into a reward that represents nobody's actual preference, and intransitive preferences cannot be represented at all. Every method fitting a Bradley-Terry reward inherits this, DPO included, since the derivation runs straight through it.
DeepWhen would you use verifiable rewards instead of a learned reward model?
Whenever correctness is decidable by a program: maths answers checked against a reference, code run against a test suite, output validated against a schema. The advantage is that the reward is not a proxy, so the usual form of reward hacking, exploiting the gap between a learned model and what it stands for, has nothing to exploit. It does not eliminate gaming entirely, since a model can special-case visible tests or find a degenerate format the checker accepts, so the pressure moves to the verifier's incompleteness. The real limit is scope: nothing verifies whether an answer was appropriately tactful, so production recipes combine verifiable rewards for reasoning and code with preference data for tone.
14Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.
Other techniques for this problem
A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.