←Home KnowML
Hands-onChapter 32

Preference Optimization: RLHF, DPO & GRPO

How a model goes from producing plausible text to producing text people want. Four methods, one objective, and a steadily shrinking amount of machinery needed to reach it.

30 min read Assumes: SFT (10), PPO and advantage (15), KL divergence (01)
Start reading
TL;DR

Supervised fine-tuning teaches format but cannot teach preference, because preference is a comparison between two acceptable answers and a label on one answer cannot express it. Everything on this page exists to learn from comparisons.

All four methods optimise the same objective: raise reward while staying close to the starting model, measured by KL divergence. They differ only in how much machinery they need. RLHF trains a reward model and runs PPO against it, holding four models in memory.

DPO proves the reward model cancels algebraically and trains on preference pairs directly. GRPO keeps the RL loop but replaces the value model with the mean reward of a sampled group. RLVR throws out the learned reward model entirely and scores against a checker that is right.

One sentence for an interview: the field spent three years discovering how much of the RLHF pipeline was removable, and the answer turned out to be most of it.

01Why supervised fine-tuning is not enough

The limitation comes from the structure of the objective, and more scale or better data quality does not remove it.

Page 10 covers SFT: show the model prompt-and-response pairs, train on next-token prediction, and it learns to answer in the right shape. That works because there is a target to imitate.

Now consider a question with two good answers. One is concise, one is thorough. One hedges appropriately, one commits. Both are fluent, both are correct, and a human asked to choose would have a clear opinion. SFT has no way to express that opinion, because its loss function only knows how to raise the likelihood of one specific token sequence.

The sentence that makes it click

SFT can only ever tell a model "produce this". It cannot tell it "produce this and not that". Preference is a relation between two outputs, and you cannot encode a relation in a per-token label on one of them.

There is a second problem, and it is the more practical one. Training on demonstrations pushes the model toward the average of its demonstrators. If the people writing your SFT data are good but not exceptional, imitation caps you at good but not exceptional.

Comparisons do not have that ceiling: people can reliably judge which of two answers is better long past the point where they could write the better one themselves.

That gap between judging quality and producing quality is what preference optimization converts into training signal.

02Intuition

Every method here is the same two-part bargain, with different amounts of machinery bolted on.

The bargain is:

  • Move toward what people prefer. Some notion of reward goes up.
  • Do not move far. Stay close to the model you started from.

The second half is not a technicality. It is the part that keeps the method from destroying the model. A policy optimising an imperfect reward with no constraint will find the exploits: repeating flattering phrases, padding length, adopting whatever surface pattern the reward signal happens to correlate with. It will score brilliantly and be useless. The distance constraint is what keeps the optimiser inside the region where the reward still means something.

The whole history of this field is a sequence of answers to one question: how much apparatus do you need to run that bargain?

Before

Instruction-tuned models were trained purely by imitation. Quality was bounded by the demonstrations, and there was no mechanism to express that one acceptable answer was better than another.

Innovation

Ouyang et al. (2022) trained a reward model on human comparisons and optimised against it with PPO. In the InstructGPT results, a 1.3B model tuned this way was preferred to a 175B model that was not.

After

Preference tuning became the default final stage of every serious model. The subsequent work has been steadily removing pieces of the pipeline while keeping the result.

03The pipeline, and where each method cuts into it

Read this as a single diagram in prose. Each method below deletes a different box.

  1. Pretraining. Next-token prediction on a large corpus. Produces a base model that can continue text but not follow instructions.
  2. SFT. It fine-tunes on demonstrations and produces a model that answers in the right shape. This checkpoint matters later, because it becomes the reference model that everything else is measured against.
  3. Reward modelling. Collect comparisons, train a model to score responses.
  4. Policy optimization. Update the SFT model to score well under that reward, penalised for drifting from the reference.

Now the cuts:

  • RLHF runs all four stages as written.
  • DPO deletes stage 3 and replaces stage 4 with a supervised loss on the preference pairs. The reward model does not just get skipped, it provably cancels.
  • GRPO keeps stages 3 and 4 but deletes the value network inside stage 4, using the average reward of several sampled answers as the baseline instead.
  • RLVR deletes stage 3 by replacing the learned reward model with a program that checks whether the answer is correct.
One pipeline, four methods — each deletes a different box
1 · PRETRAIN 2 · SFT 3 · REWARD MODEL 4 · OPTIMIZE POLICY RLHF base SFT train reward model PPO · 4 models resident DPO base SFT cancels algebraically supervised loss on pairs GRPO base SFT train reward model RL, no value model group mean is the baseline RLVR base SFT a checker, not a model RL against ground truth red strike = stage removed entirely (RLVR swaps it, DPO deletes it) faded = inherited, not retrained
Every row starts from the same SFT checkpoint, which is also the reference model the KL penalty measures against. Reading down the third column is the clearest way to see the progression: RLHF learns a reward model, DPO proves it cancels, GRPO keeps it but drops the value network beside it, and RLVR replaces it with a program that is right. Less machinery each time, same underlying objective.

04The reward model, and the assumption underneath it

A reward model turns "A is better than B" into a number. The step from comparisons to a scalar is where an assumption enters, and it is worth naming.

You cannot ask annotators to score a response out of ten. People are inconsistent at absolute judgements and drift over a session. They are far more reliable at comparisons. So the data is pairs: a prompt $x$, a preferred response $y_w$, a dispreferred one $y_l$.

To get a scalar reward out of that, you need a model of how preferences relate to scores. The standard choice is Bradley-Terry, which says the probability of preferring one item is the sigmoid of the difference in their underlying scores:

$$P(y_w \succ y_l \mid x) = \sigma\big(r(x, y_w) - r(x, y_l)\big)$$

Only the difference in rewards is identified. Adding a constant to every reward leaves every preference probability unchanged, which is why reward-model outputs have no absolute meaning and cannot be compared across training runs.

Fitting that by maximum likelihood over the dataset gives the reward model's training loss:

$$\mathcal{L}_{RM} = -\,\mathbb{E}_{(x, y_w, y_l) \sim D}\Big[\log \sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big)\Big]$$

Architecturally the reward model is usually the SFT model with the token-prediction head swapped for a single scalar output, read off the final token.

What the assumption buys, and what it costs

Bradley-Terry assumes a single consistent scalar of quality exists, and that everyone is noisily measuring the same one. This is false in an interesting way: annotators disagree about tone, caution and verbosity, and averaging their disagreement produces a reward model that represents nobody's actual preference.

It also cannot represent intransitive preferences, where A beats B, B beats C, and C beats A. Both limitations are inherited by every method on this page that fits a Bradley-Terry reward, DPO included.

05RLHF with PPO: four models in the loop

The original recipe, and still the one the others are defined against.

With a reward model in hand, the objective is the bargain from section 02, written down:

$$\max_{\pi_\theta}\ \mathbb{E}_{x \sim D,\, y \sim \pi_\theta(\cdot \mid x)}\big[r_\phi(x, y)\big] \;-\; \beta\, D_{KL}\big(\pi_\theta(\cdot \mid x)\,\|\,\pi_{\text{ref}}(\cdot \mid x)\big)$$

Maximise reward, minus $\beta$ times the KL divergence from the reference model. $\beta$ is the exchange rate between the two, and it is the single most consequential hyperparameter in the whole pipeline.

Optimising this with PPO means the RL machinery from page 15 applies directly: the language model is the policy, generating a token is an action, and the sequence so far is the state. PPO's clipped surrogate objective prevents any single update from moving the policy too far, which page 15 derives.

In practice the KL term is not applied as a separate penalty on the objective. It is folded into the reward at each token, so the per-token reward becomes the reward model's score (at the end of the sequence) minus $\beta$ times the log-ratio between policy and reference at that token. The RL algorithm then optimises a single shaped reward.

The cost of all this is the thing people complain about. You hold four models at once:

  • Policy. This is the model being trained, and it needs gradients and optimiser state.
  • Reference. A frozen copy of the SFT checkpoint, for the KL term.
  • Reward model. It is frozen and scores completions.
  • Value model. It estimates expected return from a state, which is used to compute the advantage, and it is trained alongside the policy.

Two of those need gradients and two are frozen, but all four occupy memory, and the loop also requires generating rollouts, which is slow. The next three sections exist because of this cost.

06DPO: making the reward model cancel

The result that changed practice: you do not need the reward model, and this is not an approximation.

Page 10 states the DPO loss and explains each term. What follows is why it is true, because the derivation is short and it is the difference between using DPO and understanding when it will fail.

Derivation How the reward model disappears from the RLHF objective

The claim is that optimising the RLHF objective in section 05 is equivalent to a supervised loss on preference pairs. The path runs through the closed-form solution of the KL-constrained objective.

  1. $$\max_{\pi}\ \mathbb{E}_{y \sim \pi}\big[r(x,y)\big] - \beta\, D_{KL}\big(\pi \,\|\, \pi_{\text{ref}}\big)$$
    Where we start: the RLHF objective, over all possible policies rather than a parameterised family. This is the step that makes the result exact but also marks its limit, since a real network cannot represent every distribution.
  2. $$\pi^\star(y \mid x) = \frac{1}{Z(x)}\, \pi_{\text{ref}}(y \mid x)\, \exp\!\Big(\tfrac{1}{\beta} r(x,y)\Big)$$
    Why: this objective has a known closed-form maximiser. It is the reference distribution reweighted by exponentiated reward, a Gibbs distribution, with $Z(x) = \sum_y \pi_{\text{ref}}(y \mid x)\exp(r(x,y)/\beta)$ normalising it. Computing $Z(x)$ requires summing over every possible response, so this formula cannot be used directly. That is the whole obstacle.
  3. $$r(x,y) = \beta \log \frac{\pi^\star(y \mid x)}{\pi_{\text{ref}}(y \mid x)} + \beta \log Z(x)$$
    Why: rearrange the previous line for $r$. Read it in this direction and it says something surprising: any policy implicitly defines a reward function, namely its log-ratio against the reference. The intractable $Z(x)$ is still here, but note it depends only on the prompt.
  4. $$P(y_w \succ y_l \mid x) = \sigma\big(r(x,y_w) - r(x,y_l)\big)$$
    Why: bring in Bradley-Terry from section 04, the same model the reward model was fitted with. Preferences depend only on a difference of rewards, which is what the next step exploits.
  5. $$r(x,y_w) - r(x,y_l) = \beta \log \frac{\pi^\star(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi^\star(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}$$
    The step that does the work: substitute step 3 into step 4. Both responses share the same prompt $x$, so both carry the identical $\beta \log Z(x)$ term, and subtracting removes it. The intractable normaliser is gone, and with it the need for a reward model at all.
  6. $$\mathcal{L}_{\text{DPO}} = -\,\mathbb{E}_{(x,y_w,y_l)}\left[\log \sigma\!\left(\beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}\right)\right]$$
    Where we land: maximum likelihood on the preference data under this reparameterisation, with $\pi_\theta$ in place of the unknown optimum. Every quantity is a log-probability the policy can compute directly. No reward model, no rollouts, no value network.
DPO is not an approximation of RLHF; on this objective and with an unconstrained policy class, it targets the same optimum. But the assumptions are where the failures live. It inherits Bradley-Terry, so it inherits that model's inability to represent annotator disagreement. It assumes an unconstrained policy class, which a finite network is not. And critically, it learns only from the fixed pairs in the dataset, never from its own samples, so it has no way to discover that some response it would actually produce is bad. That last point is the honest case for keeping an online method.

One consequence worth internalising: the quantity $\beta \log \frac{\pi_\theta(y \mid x)}{\pi_{\text{ref}}(y \mid x)}$ is DPO's implicit reward. You never train a reward model, but you can still read one off the policy at any point, and the gap between the implicit rewards of the chosen and rejected responses is the margin. Watching that margin is how you tell whether training is working, and section 09 is built around it.

07GRPO: dropping the value model too

DPO removes the RL loop. GRPO keeps it and removes something else, which turns out to matter enormously for reasoning.

Group Relative Policy Optimization was introduced in the DeepSeekMath paper and is the method behind DeepSeek-R1. It starts from a question about PPO: what is the value model for?

Its job is to provide a baseline. Advantage is reward minus expected reward, and subtracting a baseline is what stops every action in a good episode from being reinforced equally. PPO learns that baseline with a second network roughly the size of the policy.

GRPO's observation is that for language you can get a baseline for free. Sample a group of $G$ responses to the same prompt, score them all, and use the group's own statistics:

$$\hat{A}_i = \frac{r_i - \operatorname{mean}(r_1, \dots, r_G)}{\operatorname{std}(r_1, \dots, r_G)}$$

The advantage of response $i$ is just how much better it was than its siblings, in units of the group's own spread. No learned value function anywhere.

That advantage then goes into a PPO-style clipped objective with a KL term to the reference. The clipping and the KL constraint are unchanged; only the source of the baseline differs.

Why this fits reasoning so well

For a maths or code problem you can sample sixteen attempts and grade every one automatically. The group mean is then a good estimate of how hard this specific prompt is, which is exactly what a value model spends its capacity trying to learn.

Normalising within the group also makes easy and hard prompts contribute comparably, and easy prompts do not dominate because of higher absolute rewards. You get a per-prompt difficulty baseline for the price of sampling, and you delete a network the size of your policy.

The cost is generation: every optimisation step needs $G$ full completions per prompt and not one, so GRPO is inference-heavy in a way DPO is not. Practically this means the throughput ideas on page 31 apply directly to your training loop, and serious GRPO setups run a dedicated inference engine to produce rollouts.

The DeepSeek-R1 work pushed this further than most expected. Its central claim is that reasoning ability can be incentivised through reinforcement learning alone, without human-labelled reasoning traces to imitate. The model is rewarded for reaching correct answers and develops the intermediate reasoning by itself.

08RLVR: when the reward is a checker, not a model

The last piece to remove is the learned reward model, and for some tasks you can delete it outright.

A learned reward model is a stand-in for a human judgement you cannot query at training speed. But for a whole class of tasks there is no judgement required, only a fact:

  • A maths answer either matches the reference or it does not.
  • Code either passes the test suite or it does not.
  • Output either parses as valid JSON against a schema or it does not.

Reinforcement learning with verifiable rewards uses that check as the reward directly. The Tulu 3 work applies it as a named stage in an open post-training recipe, and it is central to how recent reasoning models are trained.

What this buys is a reward you cannot hack in the usual way. A learned reward model is an imperfect proxy, and enough optimisation pressure finds the gap between the proxy and the thing it stands for. A test suite has no gap of that kind: passing the tests is the objective.

What it does not buy is safety from all gaming. A model can still special-case the visible tests, hard-code expected outputs, or find a degenerate answer format that the checker accepts. The pressure moves from exploiting a reward model's blind spots to exploiting a verifier's incompleteness, which is a narrower target but not an empty one.

The real limit is scope. Verifiable rewards work where correctness is decidable, which excludes most of what preference tuning was invented for. Nothing checks whether an answer was appropriately tactful. In practice the two are combined: verifiable rewards for reasoning and code, preference data for tone and helpfulness.

09Build this

DPO has a pathology you will not believe until you plot it, and plotting it takes one afternoon.

Project Watch DPO push down the answer it is supposed to prefer ~4 hours · TRL + one GPU

Run DPO on a small instruction-tuned model with a public preference dataset. Then log three quantities every step: the log-probability of the chosen response, the log-probability of the rejected one, and the margin between their implicit rewards. Most tutorials plot only the margin, and the margin is the least interesting of the three.

  1. Take a small instruction-tuned model, 0.5B to 1.5B, and a standard preference dataset with prompt, chosen and rejected fields. A free Colab GPU is enough at this size with LoRA.
  2. Set up DPOTrainer from TRL with $\beta = 0.1$. With LoRA you can skip loading a separate reference model, since disabling the adapter recovers the base policy.
  3. Log logps/chosen, logps/rejected and the reward margin every step, and plot all three on shared axes.
  4. Train well past the point where the margin looks healthy, then read the two log-probability curves and not the margin.
  5. Now sweep $\beta$ with a much smaller value and a much larger one, and at each setting generate a few completions and read them.
You'll know it worked when the margin rises steadily while both log-probability curves fall. The model is becoming less likely to produce the preferred response, and DPO counts that as progress because the loss only ever sees the difference. Nothing in the objective says the chosen response should become more probable.
What the sweep teaches. Small $\beta$ weakens the tether to the reference, and generations drift toward degenerate text that scores well on the implicit reward and reads badly. Large $\beta$ holds the model so close to its starting point that the margin barely moves. You are looking directly at the bargain from section 02, and finding that neither end of the dial is where you want to be. This is also the clearest argument for why online methods exist: DPO only ever sees the dataset's pairs, so it has no way to notice that the text it now generates has become worse.

10The variants

DPO spawned a large family. Most differ in one specific way, and knowing which is enough.

IPO

Identifies that DPO's sigmoid loss can be driven arbitrarily far when a pair is perfectly separable, which overfits the preference data. Replaces it with a bounded objective that cannot run away.

Relevant when preference pairs are clean and the model overfits them quickly.

KTO

Drops the need for pairs entirely. Learns from individual responses labelled good or bad, using an objective drawn from prospect theory.

Reach for it when you have thumbs-up/thumbs-down production feedback, which is far cheaper to collect than matched pairs.

ORPO

Removes the reference model, and folds preference optimization into SFT as a single stage with an odds-ratio penalty on the rejected response.

Reach for it to halve memory and collapse two training stages into one.

SimPO

Also reference-free, using length-normalised average log-probability as the implicit reward. The normalisation directly targets DPO's tendency to reward longer responses.

Reach for it when outputs are drifting longer without getting better.

A caution applies to all of them, including the numbers in their papers: preference-tuning results are usually reported on judge-scored benchmarks, and LLM judges have documented biases, length being the best known. A method that produces longer answers can post better numbers without producing better answers. Page 25 covers why that happens. Treat published win rates as a reason to try something on your own data and not as a ranking.

11What breaks

The failure modes are shared across methods, because they come from the objective and not from the algorithm.

  • Reward hacking. The policy finds where the reward model is wrong and lives there. Classic symptom: reward climbing steadily while sampled outputs get worse. Always read generations; never trust the reward curve alone.
  • Length bias. Human annotators mildly prefer longer answers, reward models learn that preference and amplify it, and the policy exploits it. Length is the single most common confound in this entire area.
  • Both log-probabilities falling in DPO. This is the section 09 phenomenon, where the margin looks fine while the model becomes less likely to produce either response.
  • KL collapse and lost diversity. Too much optimisation pressure and the model converges on one safe phrasing for everything. Benchmarks may not notice; users do.
  • Reward model distribution shift. The reward model was trained on SFT-era outputs. As the policy improves it generates text unlike anything the reward model scored, and the scores become unreliable exactly when the policy is doing best.
  • Preference data quality. Annotator agreement on subjective comparisons is often modest, and a reward model fitted to noisy labels is a noisy reward model. This is usually the binding constraint, more than the choice of algorithm.
The habit that catches most of these

Keep a fixed set of prompts and generate from them at every checkpoint, then read the outputs yourself. Every failure above is visible in generations long before it shows up in a metric, and several of them never show up in a metric at all.

12Where you meet this in the wild

Which method fits depends almost entirely on what kind of signal you can get.

Tuning tone for a product

You have a house style and some examples of good and bad replies. Start with SFT, then DPO on a few thousand pairs. The reward model is not worth building at this scale, and DPO is stable enough to run without an RL specialist.

Training a reasoning model

Answers are checkable, so use GRPO with verifiable rewards. There is no reward model and no value model, and the signal is exactly right and not a proxy. Budget for generation: this is where most of the compute goes.

Learning from production feedback

You have thumbs-up and thumbs-down on individual responses and no matched pairs. KTO is built for this shape of data, and constructing artificial pairs from unpaired signals usually loses more than it gains.

Frontier-scale alignment

Full RLHF still appears here, because online methods can learn from the model's own current samples and not a fixed dataset. Constitutional AI adds a further step, using AI feedback against a written set of principles to reduce how much human labelling is needed.

13Interview questions

BeginnerWhy is supervised fine-tuning not enough to align a model?

Because SFT can only raise the likelihood of one target sequence, and preference is a relation between two acceptable outputs rather than a property of one. Given a concise answer and a thorough one that are both correct, SFT has no way to express that people prefer one. There is also a ceiling effect: imitation pushes the model toward the average of its demonstrators, whereas people can reliably judge which of two answers is better long past the point where they could write the better one. Preference optimization converts that judging ability into training signal.

BeginnerWhat is the KL term in the RLHF objective for?

It keeps the policy close to the SFT reference model, and it is what stops the method from destroying the model. The reward model is an imperfect proxy, so an unconstrained optimiser will find where the proxy is wrong and exploit it, producing text that scores highly and reads badly. The coefficient $\beta$ sets the exchange rate between reward and drift. Too low and the model degenerates; too high and it barely moves from where it started.

IntermediateHow does DPO eliminate the reward model, and is it an approximation?

It is not an approximation on the stated objective. The KL-constrained RLHF objective has a closed-form optimal policy, the reference distribution reweighted by exponentiated reward. Rearranging gives the reward as a log-ratio between policy and reference, plus a normalising term that depends only on the prompt. Substituting that into the Bradley-Terry preference model leaves a difference of two rewards for the same prompt, so the intractable normaliser cancels, and what remains is a supervised loss on the policy's own log-probabilities. The caveats are that the equivalence assumes an unconstrained policy class, and that DPO learns only from the fixed dataset rather than from its own samples.

IntermediateWhat does GRPO remove relative to PPO, and what does it cost?

It removes the value network. That network exists to provide a baseline for the advantage, and GRPO gets a baseline instead by sampling a group of responses to the same prompt and normalising each reward by the group's mean and standard deviation. That deletes a model roughly the size of the policy and gives a per-prompt difficulty baseline, which is exactly what the value function was trying to learn. The cost is generation: every step needs several full completions per prompt rather than one, so GRPO is inference-heavy and serious setups run a dedicated inference engine for rollouts.

IntermediateReward is climbing steadily during RLHF but the outputs look worse. What is happening?

Reward hacking. The policy has found a region where the reward model is wrong and is optimising the proxy rather than the thing it stands for. A related cause is distribution shift: the reward model was trained on SFT-era outputs, and as the policy moves away its scores become unreliable. Check whether the KL from the reference is growing, look for the usual exploits such as length inflation or repeated flattering phrasing, and read generations from a fixed prompt set at every checkpoint. Remedies are raising $\beta$, stopping earlier, or refreshing the reward model on samples from the current policy.

DeepIn DPO, both chosen and rejected log-probabilities decrease during training. Is that a bug?

It is a real and commonly observed property of the objective rather than an implementation error. The loss depends only on the difference between the two implicit rewards, so it is satisfied equally well by raising the chosen response's likelihood or by lowering the rejected one further, and gradient descent frequently does the latter to both. The consequence is that the model can become less likely to produce the preferred response while the margin looks healthy. It is a good argument for monitoring the raw log-probabilities rather than only the margin, and part of the motivation for reference-free variants and for online methods that sample from the current policy.

DeepWhat does the Bradley-Terry assumption commit you to, and when does it fail?

It assumes a single scalar of quality exists and that every annotator noisily measures the same one, with preference probability given by the sigmoid of a reward difference. Two consequences follow. Only differences are identified, so reward-model outputs have no absolute meaning and cannot be compared across runs. More seriously, genuine disagreement about tone, caution or verbosity gets averaged into a reward that represents nobody's actual preference, and intransitive preferences cannot be represented at all. Every method fitting a Bradley-Terry reward inherits this, DPO included, since the derivation runs straight through it.

DeepWhen would you use verifiable rewards instead of a learned reward model?

Whenever correctness is decidable by a program: maths answers checked against a reference, code run against a test suite, output validated against a schema. The advantage is that the reward is not a proxy, so the usual form of reward hacking, exploiting the gap between a learned model and what it stands for, has nothing to exploit. It does not eliminate gaming entirely, since a model can special-case visible tests or find a degenerate format the checker accepts, so the pressure moves to the verifier's incompleteness. The real limit is scope: nothing verifies whether an answer was appropriately tactful, so production recipes combine verifiable rewards for reasoning and code with preference data for tone.

14Go deeper

📄
Paper
Training language models to follow instructions with human feedback
Ouyang et al. — the InstructGPT paper, where the RLHF recipe was established — arXiv:2203.02155
📄
Paper
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov et al. — the derivation in section 06, in full — arXiv:2305.18290
📄
Paper
DeepSeekMath: Pushing the Limits of Mathematical Reasoning
Shao et al. — where GRPO is introduced — arXiv:2402.03300
📄
Paper
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
GRPO at scale, and the claim that reasoning can be learned without imitation data — arXiv:2501.12948
📄
Paper
Proximal Policy Optimization Algorithms
Schulman et al. — the clipped objective every method here builds on — arXiv:1707.06347
📄
Paper
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman et al. — GAE, the advantage estimator PPO uses and GRPO replaces — arXiv:1506.02438
📄
Paper
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
A fully open post-training recipe, and where verifiable rewards appear as a named stage — arXiv:2411.15124
📄
Paper
Constitutional AI: Harmlessness from AI Feedback
Bai et al. — replacing much of the human labelling with a written set of principles — arXiv:2212.08073
📄
Paper
A General Theoretical Paradigm to Understand Learning from Human Preferences
Azar et al. — the IPO paper, on how DPO overfits separable pairs — arXiv:2310.12036
📄
Paper
KTO: Model Alignment as Prospect Theoretic Optimization
Ethayarajh et al. — learning from unpaired good/bad labels — arXiv:2402.01306
📄
Paper
ORPO: Monolithic Preference Optimization without Reference Model
Hong et al. — folding preference tuning into SFT as one stage — arXiv:2403.07691
📄
Paper
SimPO: Simple Preference Optimization with a Reference-Free Reward
Meng et al. — length-normalised implicit reward — arXiv:2405.14734
🔧
Docs
TRL — DPOTrainer
The implementation for section 09, including which metrics it logs and the reference-free LoRA path.
🔧
Docs
TRL — GRPOTrainer
Group size, reward functions, and wiring a verifiable reward into the loop.
🔧
Tool
TRL
The library where every method on this page has a trainer, and the most readable reference implementations available.

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

Other techniques for this problem

A scoped slice of the full Technique Map — every technique this page covers, grouped by what it solves.

My Notes — 30 Running Models Locally

Free notes

Highlights on this page