How Models Are Trained · Part 3 · Learning From Feedback

Chapter 8 · RLHF with PPO: the full loop

RLHF with PPO from the inside: why plain REINFORCE wastes samples and takes unsafe steps, trust regions in plain words, the probability ratio and PPO's clipped objective worked by hand, the value loss and the entropy bonus, GAE on a real reply, the four models of RLHF and what they cost in memory, the implementation details that make PPO work for language models, and a from-scratch PPO run on Qwen2.5-0.5B-Instruct against the Chapter 7 reward model, checked by a stronger gold judge.

Goal: by the end of this chapter you can explain why RLHF uses PPO instead of plain REINFORCE, write down the clipped objective and compute it by hand for any token, say what the value model, the reference model and the reward model each do in the loop and what they cost, list the implementation details that decide whether a PPO run works, and read the dashboard of a real run: reward, KL, length, clip fraction, value loss. You will also have run PPO yourself on Qwen2.5-0.5B-Instruct with the reward model from Chapter 7, and checked with an independent judge how much of the reward gain was real.


8.1 Where we are, and what is still missing

Chapter 6 built reinforcement learning for text from nothing: the policy gradient, REINFORCE, baselines, the advantage, the value function, GAE, and the per-token KL penalty to a reference model. It ended with a working method, REINFORCE with a leave-one-out baseline, that turned GPT-2 into a writer of positive movie reviews, and with a list of three reasons why most RLHF systems from 2019 to 2023 used something more complicated. Chapter 7 built the judge: a reward model on Qwen2.5-0.5B-Instruct that agrees with GPT-4's preferences on 69.5% of held-out UltraFeedback pairs, likes length a little too much, and over-promises when you optimise against it.

This chapter puts the two together with the algorithm that the classic RLHF papers used: PPO, Proximal Policy Optimization (Schulman et al., 2017). Ziegler et al. (2019), Stiennon et al. (2020), InstructGPT (Ouyang et al., 2022) and Anthropic's assistant (Bai et al., 2022) all trained their policies with it. Chapter 2 told that story; here we take the method apart.

We start from what REINFORCE gets wrong for a language model. There are three problems.

  1. Each batch is used once. Generating replies is the expensive part of a step. In Chapter 6, sampling 64 replies of 24 tokens took most of each step's time; with replies of a hundred tokens or more and a bigger model, generation dominates even more. REINFORCE takes one gradient step and throws the batch away, because after the step the replies are no longer samples from the current policy, and the policy gradient formula assumes they are.
  2. Nothing limits the size of a step. The policy gradient tells us a direction. How far to go is left to the learning rate. Too small and training is slow; too large and one unlucky batch can push the policy somewhere bad, from which the next batches (sampled from that bad policy) may not rescue it.
  3. Credit per token is coarse. With one reward at the end and a baseline per reply, every token of a reply shares the same advantage, up to the KL pieces. A learned value function can do better, as Section 6.10 showed with GAE.

PPO answers all three: a probability ratio lets it reuse a batch for several gradient steps, clipping that ratio keeps each update small, and a value model with GAE gives each token its own advantage.

What each method does with one batch of 32 sampled repliesREINFORCE (Chapter 6)sample 32g1one gradient step, then the batch is thrown awayPPO (this chapter)sample 32g1g2g3g42 epochs x 2 minibatches = 4 gradient steps on the same batchSampling 32 replies of up to 128 tokens takes most of a step; reusing it is where PPO saves time.
The first difference. REINFORCE takes one gradient step per batch of samples. PPO takes several: our run in Section 8.9 makes 2 passes over each batch in 2 minibatches, so 4 gradient steps for each round of sampling.

8.2 Re-using a batch: the probability ratio

Why can't REINFORCE simply take ten steps on the same batch? Recall its loss from Section 6.5, averaged over all sampled tokens:

LPG(θ)=−1N∑i=1NA^i log⁡πθ(ai∣si)L^{\text{PG}}(\theta) = -\frac{1}{N}\sum_{i=1}^{N} \hat A_i \,\log \pi_\theta(a_i \mid s_i)

where:

  • ii runs over the NN sampled tokens in the batch (every token of every reply),
  • sis_i is the state (prompt plus the reply so far) and aia_i the token that was sampled there,
  • πθ(ai∣si)\pi_\theta(a_i \mid s_i) is the probability the current policy gives that token,
  • A^i\hat A_i is the advantage estimate for that token, a fixed number computed before the update.

Its gradient is the policy gradient only at the parameters that produced the samples. After one step the parameters have changed, and the batch is now a sample from an older policy. Keep minimising LPGL^{\text{PG}} on it and you are no longer following the gradient of anything meaningful: tokens with positive advantage get pushed up without limit, because the loss keeps rewarding a higher log⁡π\log \pi no matter how high it already is. The PPO paper says this directly:

The fix for the first half of the problem, "the samples come from an older policy", is a standard statistical tool.

Let πθold\pi_{\theta_{\text{old}}} be the policy that sampled the batch and πθ\pi_\theta the policy being updated. For each sampled token define the probability ratio:

ri(θ)=πθ(ai∣si)πθold(ai∣si)=exp⁡(log⁡πθ(ai∣si)−log⁡πθold(ai∣si))r_i(\theta) = \frac{\pi_\theta(a_i \mid s_i)}{\pi_{\theta_{\text{old}}}(a_i \mid s_i)} = \exp\big(\log \pi_\theta(a_i \mid s_i) - \log \pi_{\theta_{\text{old}}}(a_i \mid s_i)\big)

where:

  • πθold(ai∣si)\pi_{\theta_{\text{old}}}(a_i \mid s_i) is the probability of the sampled token under the policy that sampled it, stored when the batch was generated and never changed afterwards,
  • πθ(ai∣si)\pi_\theta(a_i \mid s_i) is its probability under the policy as it is now, during the update,
  • the right-hand form is how code computes it: from two log-probabilities, which is numerically safer than dividing small numbers.

At the start of the update the two policies are the same, so every ratio is exactly 1. A ratio of 1.3 means the updated policy now gives this token 30% more probability than the policy that sampled it; 0.7 means 30% less.

The surrogate objective of TRPO weights each advantage by its ratio:

LCPI(θ)=1N∑i=1Nri(θ) A^iL^{\text{CPI}}(\theta) = \frac{1}{N}\sum_{i=1}^{N} r_i(\theta)\, \hat A_i

where CPI stands for "conservative policy iteration" (Kakade and Langford, 2002), where this objective first appeared, and the other symbols are as above.

Two facts make it the right thing to optimise. First, by importance sampling, it estimates how much better than the old policy the new one would do, using only the old samples. Second, its gradient at θ=θold\theta = \theta_{\text{old}} is exactly the policy gradient. To see this, use ∇θr=r ∇θlog⁡πθ\nabla_\theta r = r \,\nabla_\theta \log \pi_\theta (the log-derivative trick of Section 6.4 again):

∇θLCPI=1N∑iA^i ri(θ) ∇θlog⁡πθ(ai∣si)  →  θ=θold, ri=1    1N∑iA^i ∇θlog⁡πθ(ai∣si)\nabla_\theta L^{\text{CPI}} = \frac{1}{N}\sum_i \hat A_i \, r_i(\theta)\, \nabla_\theta \log \pi_\theta(a_i \mid s_i) \;\xrightarrow{\;\theta = \theta_{\text{old}},\ r_i = 1\;}\; \frac{1}{N}\sum_i \hat A_i\, \nabla_\theta \log \pi_\theta(a_i \mid s_i)

where the arrow means "evaluated at the old parameters, where every ratio is 1", and the right-hand side is minus the gradient of LPGL^{\text{PG}}: the REINFORCE gradient. So the first step of optimising LCPIL^{\text{CPI}} is a REINFORCE step. Later steps differ, because the factor rir_i in front of each term records how much the policy has already moved on that token.

Worked example. A token had probability 0.100 under the old policy and advantage +1.5+1.5. After a few gradient steps the policy gives it 0.130. Its ratio is 0.130/0.100=1.300.130 / 0.100 = 1.30 and its term in LCPIL^{\text{CPI}} is 1.30×1.5=1.951.30 \times 1.5 = 1.95. The gradient of that term with respect to the token's log-probability is rA^=1.95r \hat A = 1.95: the objective still wants to push this token up, and even harder than at the start (1.5), because the ratio multiplies the advantage. Nothing in LCPIL^{\text{CPI}} says "you have moved enough". That is the second half of the problem.

8.3 Trust regions: only trust the batch near where it was sampled

A batch of 32 replies is a small, noisy sample. It tells us something reliable about the policy close to the one that produced it, and almost nothing about policies far away. TRPO (Schulman et al., 2015) turned this intuition into an algorithm: maximise the surrogate, but only inside a region around the old policy where the surrogate can be trusted.

In plain words, TRPO's rule is: improve the surrogate as much as you can, as long as the average KL divergence between the old and new policy, over the states in the batch, stays below a small number δ\delta. Solving that constrained problem needs second-order machinery (the conjugate gradient method and a line search), which is awkward for networks with billions of parameters and for code that is otherwise plain Adam. PPO's question was: can a first-order method, ordinary minibatch gradient descent, get the same protection?

Before answering, it is worth seeing the danger with numbers. ch8_toy.py builds a six-token "bandit": one state, six possible tokens with true mean rewards 0.3,0.2,0.1,0.0,−0.1,−0.20.3, 0.2, 0.1, 0.0, -0.1, -0.2, each observed with noise of standard deviation 1.0. That noise is large compared with the gaps, as it is for a reward model scoring replies. The old policy slightly prefers the middling tokens (expected reward +0.050+0.050). We sample one batch of 32 tokens, compute advantages as reward minus the batch mean, and then take 60 Adam steps on that one batch with three different objectives: the REINFORCE loss reused, the ratio objective LCPIL^{\text{CPI}}, and PPO's clipped objective (Section 8.4). Repeated for 500 different batches:

plain text
  reinforce after 60 steps: true expected reward +0.098 (10th percentile +0.010), KL(old||new) 5.222, worse than before in 22% of seeds
  ratio     after 60 steps: true expected reward +0.147 (10th percentile -0.098), KL(old||new) 3.549, worse than before in 25% of seeds
  clip      after 60 steps: true expected reward +0.063 (10th percentile +0.041), KL(old||new) 0.042, worse than before in 24% of seeds
02460204060REINFORCE, reusedratio x APPO clippedgradient steps on the same batchKL(old || new), natsHow far the policy moves-0.10-0.05+0.00+0.050204060gradient steps on the same batchtrue reward, worst 10% of seedsThe unlucky batchesstart +0.050
Sixty gradient steps on one batch of 32, averaged over 500 batches. Left: without clipping, the policy keeps moving away from the one that sampled the batch (KL above 3.5 nats); with clipping it stops near 0.04. Right: the worst 10% of batches. Without clipping they end well below where they started; with clipping they stay above it.

Read the three lines as a gamble. With no limit, the policy moves far (a KL of 3.5 to 5.2 nats means it is now nearly deterministic) towards whichever token looked best in these 32 noisy samples. On average that is a better token, so the mean goes up. But in the worst tenth of batches, the ratio objective ends at −0.098-0.098, far below the start: it bet everything on a token that only looked good by chance. The clipped objective moves about a hundred times less (KL 0.042), gains less from one batch, and its worst tenth still ends above the start.

One batch is not the whole story, though. In real training the next batch is sampled from the new policy. A policy that has collapsed onto one token samples only that token, every advantage becomes zero, and learning stops. The second experiment runs the full loop: 40 iterations, each sampling a fresh batch of 32 from the current policy and taking either 1 or 10 gradient steps on it.

plain text
  reinforce_K1  expected reward after 10/20/40 iterations: +0.097 / +0.144 / +0.215   10th percentile at 40: +0.161   P(best token) at 40: 0.52   stuck on a worse token: 0%
  reinforce_K10 expected reward after 10/20/40 iterations: +0.220 / +0.234 / +0.238   10th percentile at 40: +0.100   P(best token) at 40: 0.56   stuck on a worse token: 41%
  ratio_K10     expected reward after 10/20/40 iterations: +0.224 / +0.239 / +0.244   10th percentile at 40: +0.101   P(best token) at 40: 0.61   stuck on a worse token: 35%
  clip_K10      expected reward after 10/20/40 iterations: +0.191 / +0.243 / +0.271   10th percentile at 40: +0.199   P(best token) at 40: 0.77   stuck on a worse token: 11%
0.00.10.20.3010203040iterations (a fresh batch of 32 each)true expected reward (mean of 500 seeds)REINFORCE, 1 step per batchratio x A, 10 stepsPPO clipped, 10 stepsSeeds stuck on a worse token at the endREINFORCE, 1 step0%REINFORCE, 10 steps41%ratio x A, 10 steps35%PPO clipped, 10 steps11%"Stuck": the policy puts over 95% of itsprobability on one token that is not thebest one, so new batches teach it nothing.
The full loop over 500 seeds. One REINFORCE step per batch is safe but slow. Ten unclipped steps per batch are fast at first, then stall: about 4 in 10 seeds lock onto a worse token and never leave. Ten clipped steps end with the highest reward and the most seeds on the best token.

This is the trade PPO makes, measured:

  • One step per batch never gets stuck, but after 40 batches it has only reached +0.215+0.215.
  • Ten unclipped steps per batch use each batch more, so they are ahead after 10 iterations (+0.224+0.224 against +0.097+0.097). But 35% to 41% of the seeds end almost deterministic on a token that is not the best, and from there no batch can teach them anything.
  • Ten clipped steps per batch get most of the speed (ahead of one-step REINFORCE at every point) without most of the risk: the best final reward (+0.271+0.271), the best worst-case (+0.199+0.199), and the best token holds 77% of the probability on average.

The toy is small and its numbers belong to it, but the shape of the result is the one the PPO paper reported on robot-control tasks (Section 8.4 shows their Table 1).

8.4 The clipped objective

PPO's main objective, in the paper's notation:

LCLIP(θ)=1N∑i=1Nmin⁡(ri(θ) A^i,    clip(ri(θ), 1−ϵ, 1+ϵ) A^i)L^{\text{CLIP}}(\theta) = \frac{1}{N}\sum_{i=1}^{N} \min\Big(r_i(\theta)\,\hat A_i,\;\; \mathrm{clip}\big(r_i(\theta),\, 1-\epsilon,\, 1+\epsilon\big)\,\hat A_i\Big)

where:

  • ri(θ)r_i(\theta) is the probability ratio of token ii (Section 8.2),
  • A^i\hat A_i is its advantage, fixed during the update,
  • clip(r,1−ϵ,1+ϵ)\mathrm{clip}(r, 1-\epsilon, 1+\epsilon) cuts rr to the interval [1−ϵ,1+ϵ][1-\epsilon, 1+\epsilon]: values below 1−ϵ1-\epsilon become 1−ϵ1-\epsilon, values above 1+ϵ1+\epsilon become 1+ϵ1+\epsilon,
  • ϵ\epsilon (epsilon) is the clip range, almost always 0.2, so the interval is [0.8,1.2][0.8, 1.2],
  • min⁡\min takes the smaller of the unclipped and the clipped term, for each token separately,
  • the policy maximises LCLIPL^{\text{CLIP}}; code minimises its negative.

The best way to understand it is token by token. There are only four situations, depending on the sign of the advantage and on which way the ratio has moved.

-10120.60.81.01.21.41.6ratio r = pi_new / pi_oldobjective for this tokenadvantage A = +1.5r=1.30r=0.90-1.2-0.8-0.400.60.81.01.21.41.6ratio r = pi_new / pi_oldobjective for this tokenadvantage A = -0.8r=0.70r=1.50clipped objective (what PPO maximises)unclipped r x Athe four worked cases
One token's term in the objective as a function of its ratio. Left, a good token (A = +1.5): the objective rises with the ratio until 1.2, then goes flat, so pushing further earns nothing. Right, a bad token (A = -0.8): the objective improves as the ratio falls, until 0.8, then goes flat. On the other side of 1 the clipped and unclipped lines coincide: mistakes are always fully counted.

Worked example: the four cases. Each token had probability 0.100 under the old policy (log⁡πold=−2.303\log \pi_{\text{old}} = -2.303). After some gradient steps its probability has changed. ch8_toy.py computes each term and, with autograd, the gradient of the term with respect to the token's new log-probability:

plain text
== 1. the clipped objective for one token (epsilon = 0.2)
  A > 0, ratio went up a lot     p_old 0.100 p_new 0.130 ratio 1.301  A +1.5  r*A +1.951  clip(r)*A +1.800  min +1.800  d objective / d log p +0.000
  A > 0, ratio went down         p_old 0.100 p_new 0.090 ratio 0.900  A +1.5  r*A +1.350  clip(r)*A +1.350  min +1.350  d objective / d log p +1.350
  A < 0, ratio went down a lot   p_old 0.100 p_new 0.070 ratio 0.700  A -0.8  r*A -0.560  clip(r)*A -0.640  min -0.640  d objective / d log p +0.000
  A < 0, ratio went up           p_old 0.100 p_new 0.150 ratio 1.501  A -0.8  r*A -1.201  clip(r)*A -0.960  min -1.201  d objective / d log p -1.201
  1. Good token, already pushed up a lot. r=0.130/0.100=1.30r = 0.130 / 0.100 = 1.30. Unclipped: 1.30×1.5=1.951.30 \times 1.5 = 1.95. Clipped: 1.2×1.5=1.801.2 \times 1.5 = 1.80. The minimum is 1.80, the clipped term, which does not depend on θ\theta at all, so the gradient is 0. This token has had its share of this batch's credit; PPO stops pushing it.
  2. Good token whose probability went down (other updates in the minibatch pulled it down to 0.090). r=0.90r = 0.90, inside the interval, so both terms are 0.90×1.5=1.350.90 \times 1.5 = 1.35, and the gradient is rA^=1.35r \hat A = 1.35: keep pushing it up. Even if rr had fallen below 0.8, the minimum would pick the unclipped term (rA^r \hat A is then smaller than 0.8A^0.8 \hat A), so the gradient would stay. A good token losing probability is a mistake, and mistakes are always corrected.
  3. Bad token, already pushed down a lot. r=0.70r = 0.70. Unclipped: 0.70×(−0.8)=−0.560.70 \times (-0.8) = -0.56. Clipped: 0.8×(−0.8)=−0.640.8 \times (-0.8) = -0.64. The minimum is −0.64-0.64, the clipped, constant term: gradient 0. The bad token has been pushed down enough for one batch.
  4. Bad token whose probability went up. r=1.50r = 1.50. Unclipped: 1.50×(−0.8)=−1.201.50 \times (-0.8) = -1.20. Clipped: 1.2×(−0.8)=−0.961.2 \times (-0.8) = -0.96. The minimum is the unclipped −1.20-1.20, so the gradient is rA^=−1.20r \hat A = -1.20: push it back down, at full strength.

The pattern: the objective is clipped only on the side where the policy has already moved in the direction the advantage wanted. In the two "wrong-way" cases the minimum chooses the unclipped term and the gradient is the full rA^r \hat A. That is why the objective is a pessimistic bound: for every token it takes the less flattering of the two estimates.

The loss in code (simplified from ch8_ppo.py):

python
new_lp = token_logprobs(policy, ids, att, temp)          # log pi_theta of every sampled token, with gradient
ratio = torch.exp(new_lp - old_lp)                        # old_lp was stored when the batch was sampled
pg1 = -A * ratio                                          # minus the unclipped term
pg2 = -A * torch.clamp(ratio, 1 - 0.2, 1 + 0.2)           # minus the clipped term
pg_loss = (torch.max(pg1, pg2) * mask).sum() / mask.sum() # max of the negatives = min of the objectives

Line by line: the log-probabilities of the sampled tokens under the current policy; the ratio from the difference of logs; the two terms with a minus sign (we minimise); torch.max of the negatives, which is the minimum of the objectives; and an average over the reply tokens only (mask is 1 on tokens the policy chose and 0 on the prompt and padding).

Was the minimum worth it? The PPO paper compared it with the alternatives on seven simulated robot-control tasks:

Two different KL penalties appear in this chapter, and they are easy to confuse. The one in Table 1 is between the old and the new policy of one update: it is PPO's step-size control, the job clipping does in our run. The one from Chapter 6 is between the policy and the reference model: it keeps the whole run close to where it started, so that the reward model is not exploited. RLHF uses clipping for the first and the per-token penalty for the second.

8.5 The rest of the loss: value and entropy

PPO trains two things at once: the policy, and a value function used for the advantages. The paper writes one combined objective:

The value loss. The value model VϕV_\phi (parameters ϕ\phi) is trained by regression towards the returns computed with GAE. Most RLHF code, following the original OpenAI implementation, also clips the value update:

LVF(ϕ)=12N∑i=1Nmax⁡((Vϕ(si)−R^i)2,  (Vold(si)+clip(Vϕ(si)−Vold(si),−ϵv,ϵv)−R^i)2)L^{\text{VF}}(\phi) = \frac{1}{2N}\sum_{i=1}^{N} \max\Big(\big(V_\phi(s_i) - \hat R_i\big)^2,\; \big(V_{\text{old}}(s_i) + \mathrm{clip}(V_\phi(s_i) - V_{\text{old}}(s_i), -\epsilon_v, \epsilon_v) - \hat R_i\big)^2\Big)

where:

  • Vϕ(si)V_\phi(s_i) is the value model's current estimate for state sis_i,
  • Vold(si)V_{\text{old}}(s_i) is its estimate when the batch was collected (stored, like log⁡πold\log \pi_{\text{old}}),
  • R^i=A^i+Vold(si)\hat R_i = \hat A_i + V_{\text{old}}(s_i) is the return target from GAE (the "λ\lambda-return" of Section 6.10), computed before advantage whitening,
  • ϵv\epsilon_v is the value clip range (0.2 in our run and in the N+ paper),
  • the max⁡\max takes the larger, more pessimistic error, the same idea as the policy's min⁡\min.

Worked example. A state had Vold=0.10V_{\text{old}} = 0.10 and its return target is R^=0.90\hat R = 0.90. After a gradient step the value model says Vϕ=0.50V_\phi = 0.50. The unclipped error is (0.50−0.90)2=0.16(0.50 - 0.90)^2 = 0.16. The clipped prediction is 0.10+clip(0.40,−0.2,0.2)=0.300.10 + \mathrm{clip}(0.40, -0.2, 0.2) = 0.30, with error (0.30−0.90)2=0.36(0.30 - 0.90)^2 = 0.36. The maximum is 0.36, the clipped term, which has no gradient: the value model has moved 0.2 towards its target in this batch and must wait for the next one. Value clipping is controversial: in a large study of PPO on control tasks, Andrychowicz et al. (2021) found that this kind of value clipping "hurts the performance regardless of the clipping threshold". It is in our code because it is in the RLHF reference implementations we follow, and the fraction of clipped values is logged so we can see whether it matters.

The coefficient c1c_1. When the policy and the value function share one network, c1c_1 balances their gradients. In RLHF they are usually separate networks (Section 8.7), and then c1c_1 only rescales the value model's learning rate. We use separate optimisers and no c1c_1.

The entropy bonus. S[πθ](s)S[\pi_\theta](s) is the entropy of the policy's next-token distribution. Adding it rewards keeping options open. Control tasks sometimes need it; RLHF usually sets c2=0c_2 = 0, because the KL penalty to the reference model already keeps the distribution wide, and Chapter 6 showed (Ziegler et al.'s Table 10) that an entropy bonus alone does not prevent degenerate text. We still measure entropy: a sharp fall is a sign of collapse.

8.6 Advantages for a reply: GAE with a KL penalty at every token

Section 6.10 derived GAE and its λ\lambda trade-off on a toy. In RLHF the only new thing is what the per-token rewards look like. With a KL coefficient β\beta, Section 6.12's shaped reward is:

rt=−β(log⁡πold(yt∣st)−log⁡πref(yt∣st))+1[t=T−1]  r~(x,y)r_t = -\beta\big(\log \pi_{\text{old}}(y_t \mid s_t) - \log \pi_{\text{ref}}(y_t \mid s_t)\big) + \mathbb{1}[t = T-1]\; \tilde r(x, y)

where:

  • yty_t is reply token tt, sts_t the prompt plus the earlier reply tokens, and TT the number of reply tokens (for a finished reply the last one is <|im_end|>),
  • log⁡πold\log \pi_{\text{old}} and log⁡πref\log \pi_{\text{ref}} are the log-probabilities under the policy that sampled the reply and under the frozen reference model,
  • r~(x,y)\tilde r(x, y) is the reward model's score after normalisation and the EOS rule of Section 8.8, given at the last token only.

Then the TD errors and the GAE recursion are exactly as in Chapter 6, with γ=1\gamma = 1 and λ=0.95\lambda = 0.95:

δt=rt+V(st+1)−V(st),A^t=δt+λA^t+1,R^t=A^t+V(st)\delta_t = r_t + V(s_{t+1}) - V(s_t), \qquad \hat A_t = \delta_t + \lambda \hat A_{t+1}, \qquad \hat R_t = \hat A_t + V(s_t)

where V(sT)=0V(s_T) = 0 after the last token (the episode is over), the recursion runs backwards from the last token, and R^t\hat R_t is the value target.

Worked example on a real reply. In the first batch of the β=0.05\beta = 0.05 run, one prompt asked: Determine the topic of the question. Question: "what is the source of geothermal energy?" Topic: The policy answered "The topic of this question is Geothermal Energy." in 11 tokens, counting the closing <|im_end|>. ch8_ppo.py saved every number for this reply:

plain text
 t  token          log pi_old  log pi_ref   V(s_t)     r_t    A_t (GAE)  return
 0  'The'            -0.095      -0.095     -1.652   +0.000    +0.350    -1.302
 1  ' topic'         -0.000      -0.000     -7.545   +0.000    +6.572    -0.973
 3  ' this'          -6.890      -6.890     -4.920   +0.000    +4.262    -0.658
 6  ' Ge'            -0.896      -0.896     -4.706   +0.000    +4.606    -0.100
 8  ' Energy'        -0.011      -0.011     -6.637   +0.000    +7.155    +0.518
 9  '.'              -0.017      -0.017     -3.944   +0.000    +4.696    +0.753
10  '<|im_end|>'     -0.019      -0.019     -0.964   +0.843    +1.807    +0.843

(Selected rows; the full table is in results/ch8_ppo_beta005.json.)

One reply from the first batch, every token (beta = 0.05, gamma = 1, lambda = 0.95)tokenThetopicofthisquestionisGeothermalEnergy.im_endV(s_t)-1.65-7.55-2.92-4.92-4.95-2.37-4.71-5.04-6.64-3.94-0.96r_t+0.00+0.00+0.00+0.00+0.00+0.00+0.00+0.00+0.00+0.00+0.84A_t (GAE)+0.35+6.57+2.05+4.26+4.52+2.04+4.61+5.20+7.15+4.70+1.81return-1.30-0.97-0.87-0.66-0.43-0.33-0.10+0.16+0.52+0.75+0.84Reward model score (normalised) at the last token: +0.843; the other rewards are -0.05 x (log pi_old - log pi_ref).
The same reply, all 11 tokens. In the first iteration the policy equals the reference, so every KL reward is 0 and the only reward is the score at the last token. The values come from a value model that is still a copy of the reward model, and they are far too low everywhere except at the end.
  • The rewards. At the first iteration the policy is the reference, so log⁡πold=log⁡πref\log \pi_{\text{old}} = \log \pi_{\text{ref}} on every token and every KL reward is exactly 0. The reward model gives the reply a raw score of 6.037; normalised, (6.037−4.496)/1.828=+0.843(6.037 - 4.496) / 1.828 = +0.843, which lands on the last token. (The token " this" has log⁡π=−6.890\log \pi = -6.890, a probability of 0.001: sampling at temperature 0.7 still picks unlikely tokens now and then.)
  • The last token. δ10=r10+V(s11)−V(s10)=0.843+0−(−0.964)=1.807\delta_{10} = r_{10} + V(s_{11}) - V(s_{10}) = 0.843 + 0 - (-0.964) = 1.807. Nothing follows, so A^10=1.807\hat A_{10} = 1.807 and the return is A^10+V(s10)=1.807−0.964=0.843\hat A_{10} + V(s_{10}) = 1.807 - 0.964 = 0.843: the score itself.
  • One step back. δ9=0+V(s10)−V(s9)=−0.964−(−3.944)=2.980\delta_9 = 0 + V(s_{10}) - V(s_9) = -0.964 - (-3.944) = 2.980, and A^9=2.980+0.95×1.807=2.980+1.717=4.697\hat A_9 = 2.980 + 0.95 \times 1.807 = 2.980 + 1.717 = 4.697, which the script prints as 4.696 from unrounded values.
  • Two steps back. δ8=−3.944−(−6.637)=2.693\delta_8 = -3.944 - (-6.637) = 2.693 and A^8=2.693+0.95×4.696=7.154\hat A_8 = 2.693 + 0.95 \times 4.696 = 7.154.

Look at the values: −7.5-7.5 for the state after "The", −6.6-6.6 before "Energy", and close to the real score only at <|im_end|>. The value model is still a copy of the reward model, and a reward model was trained only to read the last token; at every other position its head produces numbers with no meaning. The N+ paper shows the same picture for its reward model (their Detail 13: "most values in the reward logits are non-valid and negative"). As a result every advantage in this reply is large and positive. That is less harmful than it looks, because advantages are whitened over the whole batch: what matters is how this reply's advantages compare with those of the other 31 replies, which suffer from the same offset. Within a few iterations the value model learns real per-token values (Section 8.10 shows its explained variance rising), and from then on the advantages mean what Section 6.8 said they should.

8.7 RLHF as a loop: four models

Put the pieces together and one PPO iteration for a language model looks like this:

1 rolloutpolicy samples one reply per prompt32 prompts, T = 0.7, <= 128 tokens2 scorereward model scores each reply;policy, reference, value read every token3 rewards-beta x (log pi_old - log pi_ref) per token,plus the normalised score at the last token4 advantagesGAE (gamma 1, lambda 0.95), then whiten;returns = advantages + values5 updateclipped policy loss + clipped value loss,2 epochs x 2 minibatches on this batchrepeatno gradientsteps 1 to 4store for every token:old log-prob, ref log-prob,value, reward, advantagegradientpolicy LoRA + value model
One iteration of RLHF with PPO, as implemented in ch8_ppo.py. Steps 1 to 4 need no gradients: they produce, for every reply token, its old log-probability, its reference log-probability, its value, its reward and its advantage. Step 5 is the only one that trains, and it is repeated for several epochs on the same batch.

Four networks take part. Stiennon et al. describe the setup that everyone after them copied:

PolicyQwen2.5-0.5B-Instruct+ LoRA adapter (trains)writes replies; gets the PPO gradientReferencesame weights, adapter off(frozen)log pi_ref for the KL penaltyReward modelChapter 7 RM(frozen)one score per finished replyValue modelcopy of the Chapter 7 RMadapter + head traina value at every token, for GAEone copy in memoryvalue starts as a copy of the reward modelWhat each one does:
The four models. Two are trained (policy and value), two are frozen (reference and reward). In our run the policy and the reference share one copy of Qwen2.5-0.5B-Instruct (the reference is the policy with its LoRA adapter switched off), and the value model starts as a copy of the Chapter 7 reward model.
  • Policy πθ\pi_\theta: the model being trained, initialised from the SFT model. It generates the replies and receives the PPO gradient.
  • Reference πref\pi_{\text{ref}}: a frozen copy of the starting policy. It is only used to compute log⁡πref\log \pi_{\text{ref}} for the KL penalty.
  • Reward model rψr_\psi: frozen. It scores each finished reply once.
  • Value model VϕV_\phi: trained by the value loss. Initialised from the reward model, because the reward model already "knows" what a good reply looks like; at the start its per-token values are not good predictions of the return, but it learns them faster than a value model starting from scratch.

InstructGPT used the same four, at a larger scale: a 6B reward model and a 6B value model for policies of all sizes.

PPO-ptx. InstructGPT noticed that RLHF made the model worse on some public NLP benchmarks (SQuAD, DROP), a cost they called the "alignment tax". Their fix was to mix ordinary next-token prediction on pretraining data into the RL loss: the objective of Chapter 2's Equation 2, reward minus β\beta times KL plus γ\gamma times the log-likelihood of pretraining text. With the 27.8 coefficient and 8 pretraining examples per RL episode, the language-modelling gradient is a substantial part of every step. Our run does not do this; it is a way to protect abilities that the reward model does not measure.

Adaptive KL. Ziegler et al.'s controller (Section 6.12 worked an example) adjusts β\beta to hit a target KL instead of fixing it. InstructGPT and the N+ reproduction used a fixed β\beta; libraries offer both. We use fixed values, so that the two runs differ in one number only.

What four models cost

RLHF is expensive in memory as much as in compute. A rough count for full fine-tuning with Adam in mixed precision: a trained model needs about 16 bytes per parameter (2 for the bf16 weights, 2 for the gradients, 4 for an fp32 master copy of the weights, 4 and 4 for Adam's two moving averages); a frozen model in bf16 needs 2. That ignores activations and the cache used while generating, which add a lot for long replies.

plain text
== 3. memory for the four models (weights and optimizer state only; no activations, no KV cache)
  0.5B (our Qwen2.5-0.5B)  policy     7.9 GB  value     7.9 GB  reference    1.0 GB  reward    1.0 GB  total    17.8 GB  (SFT of the same model: 7.9 GB, so RLHF needs 2.25x)
  7B                       policy   112.0 GB  value   112.0 GB  reference   14.0 GB  reward   14.0 GB  total   252.0 GB  (SFT of the same model: 112.0 GB, so RLHF needs 2.25x)
  70B                      policy  1120.0 GB  value  1120.0 GB  reference  140.0 GB  reward  140.0 GB  total  2520.0 GB  (SFT of the same model: 1120.0 GB, so RLHF needs 2.25x)
GB of weights + optimizer state, full fine-tuning, mixed precision0.5B (our Qwen2.5-0.5B)policy: 7.9 GBvalue: 7.9 GBref: 1.0 GBrm: 1.0 GB17.8 GBpolicy 7.9 + value 7.9 + reference 1.0 + reward 1.07Bpolicy: 112.0 GBvalue: 112.0 GBref: 14.0 GBrm: 14.0 GB252.0 GBpolicy 112.0 + value 112.0 + reference 14.0 + reward 14.070Bpolicy: 1120.0 GBvalue: 1120.0 GBref: 140.0 GBrm: 140.0 GB2,520.0 GBpolicy 1,120.0 + value 1,120.0 + reference 140.0 + reward 140.0policyvalue modelreferencereward modelEach bar is scaled to its own total; the totals grow 10x per row.
Weights and optimizer state for the four models with full fine-tuning, from ch8_toy.py. The two trained models dominate. RLHF with a same-size value model needs 2.25 times the memory of supervised fine-tuning, before counting activations and generation.

Worked example, 7B. Policy: 16×7×109=11216 \times 7 \times 10^9 = 112 GB. Value model, same size: 112 GB. Reference and reward model: 2×7×109=142 \times 7 \times 10^9 = 14 GB each. Total 112+112+14+14=252112 + 112 + 14 + 14 = 252 GB, against 112 GB to fine-tune the same model with SFT: 252/112=2.25252 / 112 = 2.25 times. That is more than three 80 GB GPUs before a single activation is stored, which is why RLHF code is full of tricks: sharing a backbone between policy and value (Stiennon et al. found separate networks work better), offloading the frozen models, LoRA, and the critic-free methods that drop the value model altogether.

We fit the four models of a 0.5B policy into far less with two tricks. The policy is LoRA (rank 16 on all seven linear layers of every block, 8.8 million trainable parameters), so the reference is free: it is the same network with the adapter switched off (with policy.disable_adapter():). The value model is a second copy of the Chapter 7 reward model whose LoRA adapter and score head are trained. Everything runs in fp32 on an Apple M5 Pro (64 GB); after the first iteration the process held 39.4 GB of GPU memory, most of it activations and the 151,936-wide logits of 16 sequences at a time, not weights.

8.8 The details that decide whether it works

The PPO paper gives an algorithm; making it work for language models took years of folklore. Huang et al. (2024) reproduced OpenAI's summarisation results (Stiennon et al., 2020) from scratch and wrote down every detail they found that mattered, more than 20 of them. Their hyperparameters are close to Stiennon's:

The details that matter most for RLHF, with what our code does for each:

1. Normalise the reward model's score before RL. A Bradley-Terry reward model's scale and offset are arbitrary (Section 7.4). InstructGPT and the N+ paper shift it so that the reference answers score 0 on average. We take 192 replies of the starting policy to training prompts, score them, and use their mean and standard deviation: every score is turned into (r−μ)/σ(r - \mu) / \sigma, so 0 means "as good as an average reply of the starting model" and 1 means one standard deviation better. This also makes β\beta meaningful: a KL of 10 nats at β=0.05\beta = 0.05 costs half a standard deviation of reward.

2. Initialise the value model from the reward model.

3. The EOS trick: a fixed penalty for replies that never end. A reward model is trained to read its score at the token that closes the answer (Section 7.5). A reply cut off by the length limit has no such token, so its score is not well defined.

We give every reply that has not emitted <|im_end|> within 128 tokens a normalised score of −1-1: one standard deviation below the starting model's average, whatever the reward model would have said.

4. Sample with a temperature, and use the same temperature in the loss. N+ samples at temperature 0.7 (Table 7). A detail that is easy to miss: the logits must then be divided by 0.7 also when computing log-probabilities for the ratio and the KL, as the reference implementations do. Otherwise πold\pi_{\text{old}} is not the distribution the tokens actually came from, and the ratio at the first gradient step is not 1. We do the same.

5. Turn off dropout.

6. Whiten the advantages. After GAE, subtract the batch mean of the advantages and divide by their standard deviation, over all reply tokens (Detail 25 of N+). This keeps the size of the policy gradient stable as the reward scale drifts during training. We do it.

7. Whitening the rewards is optional. Ziegler et al.'s code could also whiten the per-token rewards before GAE (Detail 24). N+ treats it as optional; we do not use it.

8. Average the loss over tokens, mask the prompt. The loss is averaged over every reply token in the minibatch; prompt and padding tokens are masked out of the policy loss, the value loss and the KL. Implementations differ here (per-token mean, per-reply mean, or a constant divisor), and the choice changes how much each token of a long reply counts.

9. Clip the gradient norm, use Adam with a small epsilon. Gradients are clipped to norm 1 for both models, Adam's ϵ\epsilon is 10−510^{-5} as in N+, and there is no weight decay.

10. Reshuffle prompts when they run out. Training draws prompts without replacement and reshuffles when they are used up (Detail 21).

The same list, for PPO in general and not only RLHF, was compiled by Huang et al. (2022) in "The 37 Implementation Details of Proximal Policy Optimization" (see References), and Engstrom et al. (2020) showed that on control tasks such "code-level optimizations" explain much of PPO's advantage over TRPO. For RLHF, Zheng et al. (2023) ran ablations of a similar list on 7B models and named their combination PPO-max. The common message: most of these details are small on their own, and a run that ignores several of them often fails in ways that look like a problem with the method.

8.9 Hands-on: PPO on Qwen2.5-0.5B-Instruct

Everything above now goes into one script, ch8_ppo.py: about 230 lines of plain PyTorch on top of the helpers in ch8_common.py, no RL library. It trains Qwen2.5-0.5B-Instruct, the same model Chapter 7 sampled for best-of-n, against the Chapter 7 reward model.

The prompts. UltraFeedback prompts of at most 128 tokens, none of which the reward model was trained or tested on. Qwen2.5-0.5B-Instruct likes to write long answers (in Chapter 7, only 27.5% of its samples ended within 256 tokens), and long replies make every step slow. So ch8_prompts.py caps replies at 128 tokens and keeps the prompts for which one sample of the starting model ended within that cap: 258 of 1,200 training candidates and 39 of 200 test candidates. This choice favours prompts with short answers (the kept replies averaged 45 tokens); it makes the runs fast and makes "did the reply end?" a real constraint the policy can break.

The settings, following Table 7 of the N+ paper except where noted:

settingvaluenote
policyQwen2.5-0.5B-Instruct + LoRA, rank 16, alpha 32, all linear layers8.8 million trainable parameters
reward modelChapter 7, frozenscores normalised with mean 4.496, sd 1.828 of 192 starting replies
value modelcopy of the reward model, LoRA + head trainedvalues on the same normalised scale
replies32 per iteration, one per prompt, temperature 0.7, at most 128 tokensN+: 512 per batch
EOS rulenormalised score −1-1 if no `<im_end
KLper-token k1k_1 penalty, β=0.05\beta = 0.05 or β=0\beta = 0N+: 0.05
GAEγ=1\gamma = 1, λ=0.95\lambda = 0.95, advantages whitenedas N+
PPO2 epochs x 2 minibatches of 16, clip 0.2, value clip 0.2N+: 4 epochs x 1 minibatch
optimiserAdamW, learning rate 3×10−53 \times 10^{-5} for both models, ϵ=10−5\epsilon = 10^{-5}, gradient norm 1N+: 3×10−63 \times 10^{-6} for full fine-tuning; LoRA needs a larger rate
length20 iterations, 640 replies per run, one seedN+: a million episodes

The learning rate came from one short trial: at 5×10−55 \times 10^{-5} with 4 epochs, the first iteration already had an approximate KL between old and new policy of 0.098 per token and a clip fraction of 0.22, far above the usual range (Section 8.11), so we lowered the rate and the number of epochs. An iteration of the β=0.05\beta = 0.05 run took 30 to 75 seconds on the M5 Pro, of which generation was only 6 to 9 seconds: the forward and backward passes, with a softmax over the 151,936-entry vocabulary at every token, dominate. That is the main reason the runs are short.

The code

The heart of ch8_ppo.py, simplified (the full script also logs, saves checkpoints and handles memory):

python
S = generate(policy, prompts, max_new=128, temp=0.7)              # 1. rollout: 32 replies
ids, att, act = batch_tensors(S)                                   #    act = 1 on reply tokens (incl. <|im_end|>)
with torch.no_grad():                                              # 2. score, no gradients
    old_lp = token_logprobs(policy, ids, att, temp)                #    log pi_old, logits / 0.7
    with policy.disable_adapter():
        ref_lp = token_logprobs(policy, ids, att, temp)            #    log pi_ref: same weights, adapter off
    old_v = values_of(ids, att)                                    #    value at every token, normalised
    raw = score_ids(rm, rm_texts(S))                               #    reward model, one score per reply
    score = (raw - MU) / SD
    score[~finished] = -1.0                                        #    the EOS trick
    rew = -beta * (old_lp - ref_lp) * act                          # 3. per-token KL penalty ...
    rew[rows, last_token] += score                                 #    ... plus the score at the last token
    adv, ret = gae(rew, old_v * act, act, gamma=1.0, lam=0.95)     # 4. advantages and value targets
    adv = whiten(adv, act)
for epoch in range(2):                                             # 5. update, 2 epochs x 2 minibatches
    for mb in minibatches(32, size=16):
        ratio = torch.exp(token_logprobs(policy, ids[mb], att[mb], temp) - old_lp[mb])
        pg = torch.max(-adv[mb] * ratio, -adv[mb] * ratio.clamp(0.8, 1.2))
        v = values_of(ids[mb], att[mb])
        v_clip = old_v[mb] + (v - old_v[mb]).clamp(-0.2, 0.2)
        vl = 0.5 * torch.max((v - ret[mb]) ** 2, (v_clip - ret[mb]) ** 2)
        loss = masked_mean(pg, act[mb]) + masked_mean(vl, act[mb])
        loss.backward(); clip_grad_norm_(..., 1.0); opt_p.step(); opt_v.step()

Block by block:

  1. Rollout. One reply per prompt, sampled with Hugging Face's generate from the current policy. batch_tensors builds prompt-plus-reply sequences (with the closing <|im_end|> for replies that finished) and a mask act that is 1 exactly on the tokens the policy chose.
  2. Score. Three forward passes without gradients: the policy (its log-probabilities become log⁡πold\log \pi_{\text{old}}, fixed for this iteration), the same network with the adapter disabled (log⁡πref\log \pi_{\text{ref}}), and the value model. values_of runs the value model's transformer and applies its score head to every position, not only the last, which is what turns a reward model into a value model. Then the reward model scores each reply as in Chapter 7, the score is normalised, and replies that never ended get −1-1.
  3. Rewards. The per-token reward of Section 8.6: −β-\beta times the log-ratio on every reply token, plus the score on the reply's last token.
  4. Advantages. The backward GAE recursion of Section 6.10 (gae handles right-padded rows: the value after a reply's last token is 0), then whitening over all reply tokens of the batch.
  5. Update. The clipped policy loss of Section 8.4 and the clipped value loss of Section 8.5, averaged over reply tokens, two passes over the batch in minibatches of 16. The policy's LoRA adapter and the value model have separate AdamW optimisers.

The value model needs one more line of explanation. In Chapter 7 the reward model read the hidden state of the last token. Here, values_of reads every position:

python
def values_of(ids, att):
    inner = value.base_model.model                     # Qwen2ForSequenceClassification with LoRA layers inside
    h = inner.model(input_ids=ids, attention_mask=att).last_hidden_state
    v = inner.score(h)[..., 0].float()                 # the 896 -> 1 head, applied at every position
    return ((v - MU) / SD)[:, :-1]                     # position t = the state before choosing token t+1

The hidden state at position tt has seen the prompt and the reply up to token tt, which is exactly the state in which the policy chose token t+1t+1. Hence the shift by one at the end, the same shift that aligns logits with the tokens they predict (Chapter 1).

What one iteration measures

Every iteration logs a line like this one (the first iteration of the β=0.05\beta = 0.05 run):

plain text
step   0  score -0.125  rm_raw +4.178  kl   0.00  len  71.6  fin 0.66  ent 0.830  clipfrac 0.062  approx_kl 0.0132  v_loss 4.057  ev -0.55  (54s, gen 6s)
  • score: mean normalised reward model score of the 32 replies, after the EOS rule. The starting model's replies average about 0 by construction; this batch is slightly below.
  • rm_raw: the same scores before normalisation, on the reward model's own scale.
  • kl: the sum of log⁡πold−log⁡πref\log \pi_{\text{old}} - \log \pi_{\text{ref}} over each reply's tokens, averaged over replies: the k1k_1 estimate of Section 6.12, in nats per reply. Zero at the start, since the policy is the reference.
  • len and fin: mean reply length in tokens and the share of replies that ended.
  • ent: the policy's mean entropy per reply token, in nats.
  • clipfrac: during the update, the share of reply tokens whose ratio was outside [0.8,1.2][0.8, 1.2].
  • approx_kl: during the update, the mean over tokens of (r−1)−log⁡r(r - 1) - \log r, Schulman's k3k_3 estimate of KL(πold∥πθ)\mathrm{KL}(\pi_{\text{old}} \Vert \pi_\theta): how far each update moved the policy from the one that sampled the batch.
  • v_loss and ev: the value loss and the explained variance of the returns by the old values. An evev of −0.55-0.55 means the freshly copied reward model is a worse predictor of the returns than their plain average, as expected for a reward model read at tokens it was never trained on.

8.10 Results

Two runs, identical except for β\beta: the same seed, the same prompts in the same order, the same starting adapter. One uses the N+ value β=0.05\beta = 0.05, the other no KL penalty at all. Each ran for 20 iterations (640 replies). One seed per run, so differences smaller than the batch-to-batch noise should not be read as real.

What the training batches show

-0.2-0.100.10.205101520normalised proxy score (RM, EOS rule)024605101520KL to the reference, nats per reply2040608005101520PPO iterationreply length, tokens0.60.70.80.9105101520PPO iterationshare of replies that endedbeta = 0.05no KL penalty5-iteration running mean of each batch
Training batches of the two runs, 5-iteration running means. The proxy score rises a little and noisily. The clearest changes are elsewhere: replies get shorter (from about 53 to 31 tokens) and almost all of them end, so fewer collect the -1 of the EOS rule. Without a KL penalty the policy drifts further from the reference.

Averages over the first five and the last five iterations of each run, from the hist entries in results/ch8_ppo_<run>.json:

β=0.05\beta = 0.05no KL
proxy score (EOS rule), iterations 0 to 4−0.018-0.018−0.012-0.012
proxy score (EOS rule), iterations 15 to 19+0.095+0.095+0.129+0.129
reply length, tokens52.6 to 31.052.9 to 30.9
share of replies that end0.81 to 0.910.81 to 0.92
KL to reference, iterations 15 to 19 (nats per reply)3.495.26
entropy per token (nats)0.817 to 0.7430.822 to 0.733

Three things happened, and only one of them is what you might expect from "RLHF makes answers better".

  1. The policy learned to finish. The share of replies that end within 128 tokens went from about 81% to about 91%, and the average reply got 40% shorter. With the EOS rule, a reply that is cut off scores −1-1 whatever it says, while the starting model's finished replies average about 0. The cheapest way to raise the reward was to stop getting cut off. This is the "constraining completion length" effect of the N+ paper, and in our setup it outweighs the reward model's own slight preference for length.
  2. The proxy score rose, a little. About +0.1+0.1 standard deviations over 20 iterations. The batch-to-batch noise (32 different prompts each time) is about as large, so the curve is jagged.
  3. Without the penalty, the policy moved further. Both runs started at the same place and saw the same prompts; by the end the unpenalised policy was at 5.3 nats from the reference against 3.5 for β=0.05\beta = 0.05. At β=0.05\beta = 0.05, 3.5 nats cost 0.05×3.5=0.170.05 \times 3.5 = 0.17 in reward, comparable to the reward gained, which is why the penalty bites even this early.

Is it real? The gold judge

The training reward is measured on the training prompts by the reward model being optimised. ch8_eval.py checks the checkpoints after 10 and 20 iterations on the 39 held-out prompts, two replies each (78 replies per checkpoint, the same prompts and the same sampling seed for every checkpoint), with two judges: our reward model (the proxy) and Skywork-Reward-V2-Qwen3-0.6B (the gold), the stronger reward model Chapter 7 used. Both scores are in standard-deviation units of the starting model. ch8_analysis.py pairs each reply with the starting model's reply to the same prompt and gives 95% bootstrap intervals:

plain text
78 held-out replies per checkpoint (same prompts, same sampling seed)
beta005  step 10: proxy change -0.008 [-0.124, +0.108]  gold change +0.129 [-0.076, +0.333]  proxy with EOS rule +0.140 (start +0.071)  finished 0.94 (start 0.79)  tokens 30.9 (start 59.3)  KL 2.69
beta005  step 20: proxy change +0.074 [-0.053, +0.197]  gold change -0.044 [-0.232, +0.172]  proxy with EOS rule +0.203 (start +0.071)  finished 0.90 (start 0.79)  tokens 35.9 (start 59.3)  KL 3.20
nokl     step 10: proxy change +0.012 [-0.104, +0.132]  gold change +0.113 [-0.085, +0.313]  proxy with EOS rule +0.185 (start +0.071)  finished 0.94 (start 0.79)  tokens 30.7 (start 59.3)  KL 2.86
nokl     step 20: proxy change +0.018 [-0.103, +0.131]  gold change +0.042 [-0.138, +0.230]  proxy with EOS rule +0.185 (start +0.071)  finished 0.94 (start 0.79)  tokens 28.8 (start 59.3)  KL 5.65
-0.20+0.00+0.20+0.4001020proxy (our RM): 0, +0.20proxy (our RM): 10, +0.19proxy (our RM): 20, +0.27gold (Skywork): 0, -0.00gold (Skywork): 10, +0.13gold (Skywork): 20, -0.04PPO iteration (checkpoint)beta = 0.05-0.10+0.00+0.10+0.20+0.3001020proxy (our RM): 0, +0.20proxy (our RM): 10, +0.21proxy (our RM): 20, +0.21gold (Skywork): 0, -0.00gold (Skywork): 10, +0.11gold (Skywork): 20, +0.04PPO iteration (checkpoint)no KL penaltyproxy: our Chapter 7 reward model, raw score (no EOS rule)gold: Skywork-Reward-V2-Qwen3-0.6B (never seen in training)both in sd units of the starting model
Held-out replies at the start and after 10 and 20 iterations. Dashed: our reward model (no EOS rule). Solid: the gold reward model. Changes are small and every interval includes zero; the gold peaks at iteration 10 in both runs and falls back by iteration 20.
-0.20+0.00+0.20+0.400123456beta = 0.05, gold: 0, -0.00beta = 0.05, gold: 3, +0.13beta = 0.05, gold: 3, -0.04beta = 0.05, proxy: 0, +0.20beta = 0.05, proxy: 3, +0.19beta = 0.05, proxy: 3, +0.27no KL penalty, gold: 0, -0.00no KL penalty, gold: 3, +0.11no KL penalty, gold: 6, +0.04no KL penalty, proxy: 0, +0.20no KL penalty, proxy: 3, +0.21no KL penalty, proxy: 6, +0.21beta = 0.05, proxyno KL penalty, proxyno KL penalty, goldbeta = 0.05, goldKL to the starting model on held-out replies (nats per reply)reward (sd units)
The same checkpoints plotted against their KL from the starting model, Gao-style. Solid: gold; dashed: proxy. The no-KL run reaches iteration 20 at almost twice the KL of the penalised run without a better gold score.

Read honestly, the held-out numbers say:

  • The objective PPO optimised did go up. The proxy with the EOS rule, which is what the policy is paid, rose from +0.071+0.071 to between +0.14+0.14 and +0.20+0.20 on prompts it never saw, mostly because the share of finished replies rose from 0.79 to 0.90 to 0.94.
  • The reward model's own score barely moved. Without the EOS rule the proxy changed by −0.008-0.008 to +0.074+0.074, all within the noise.
  • The gold went up, then came back. At iteration 10 both runs are about +0.12+0.12 above the start on gold, larger than their proxy change; shorter, finished answers are something Skywork likes. By iteration 20 the gold is back near zero (−0.044-0.044 and +0.042+0.042), while the proxy with the EOS rule held or rose. No interval excludes zero, so this is a hint, not a result. But it is the shape Gao et al. describe: gold first rises with KL, then falls while the proxy keeps going.
  • The penalty bought distance, not quality. At iteration 20 the unpenalised policy is 5.65 nats from the start, the penalised one 3.20, with no measurable difference in gold. That is Gao et al.'s "akin to early stopping" result in miniature: the penalty slows the drift; it does not make the reward model more trustworthy.

What we did not see in 20 iterations: no length explosion (lengths fell), no collapse into repeated phrases (distinct-2 stayed between 0.88 and 0.90 at every checkpoint), no entropy crash (0.82 to 0.74 nats per token). With a reward model as weak as ours, longer runs would very likely find its holes, as best-of-n did in Chapter 7; our budget stopped well before that.

Before and after

The same held-out prompts, the same sampling seed, replies from the starting model and from both runs after 20 iterations (from results/ch8_analysis_stdout.txt; scores are proxy and gold in starting-model standard deviations):

plain text
PROMPT: 'Act as a Salesforce administrator expert. If okay, just say "Okay".'
  [start    ] proxy -0.47 gold -0.18 tokens  18 "Okay. I'm here to help! What would you like to know about Salesforce administrators?"
  [beta 0.05] proxy +0.52 gold +0.54 tokens   2 'Okay.'
  [no KL    ] proxy +0.52 gold +0.54 tokens   2 'Okay.'

PROMPT: 'Detailed Instructions: ... classify given tweet text into two categories: 1) positive, and 2) negative ...'
  [start    ] proxy +1.18 gold +1.32 tokens  45 'Based on the content of the tweet provided, it appears to be expressing excitement about moving to California for summer vacation, which is likely seen as a positive experience. Therefore, I would classify this tweet as having a positive classification.'
  [beta 0.05] proxy +1.28 gold +0.38 tokens  13 'Based on the content provided, this tweet appears to be:\nPositive'
  [no KL    ] proxy +0.30 gold -1.18 tokens  28 'Based on the provided tweet text, this appears to be an informative and enthusiastic statement about living in Phoenix for the time being moving to California.'

PROMPT: 'Could you paraphrase the instruction for adding potatoes and mushrooms? ...'
  [start    ] proxy +0.22 gold -0.29 tokens 111 "Sure! Here's a paraphrased version of the instruction: ..."
  [beta 0.05] proxy +0.36 gold -0.04 tokens  77 'For the mashed potato and mushroom dish, combine all remaining ingredients except potatoes and mushrooms ...'
  [no KL    ] proxy +0.10 gold -1.26 tokens  72 'For this recipe, follow these steps:\n1. Add potatoes and mushrooms to an ovenproof quart dish. ...'
  • "Just say Okay." The starting model said more than it was asked; both trained policies say exactly "Okay.", and both judges approve. This is the good side of learning to be brief.
  • The tweet. The penalised policy answers with the label, as asked. The unpenalised one has drifted into describing the tweet without classifying it; the gold model drops it from +1.32+1.32 to −1.18-1.18, our reward model only to +0.30+0.30.
  • The recipe. The unpenalised reply tells you to add the potatoes and mushrooms first, the opposite of the instruction; the gold judge marks it down to −1.26-1.26, our reward model barely notices.

Three prompts prove nothing, but they show what the averages hide: the same small average change is made of real improvements and real regressions, and the weaker judge misses more of the regressions.

8.11 What breaks, and what to watch

PPO for language models fails in a handful of recognisable ways. Each has a signature on the dashboard.

Reward hacking. The policy finds what the reward model pays for that people would not. Chapter 6 saw it with a sentiment classifier ("beautifully captures this superb masterpiece"), Chapter 7 measured its best-of-n version, and Gao et al. (2022) mapped it for PPO: the gold reward first rises with KL, then falls, while the proxy keeps rising. Their result on the KL penalty is worth knowing because it is not what most people expect:

Length exploitation. Reward models tend to prefer longer answers (Section 7.9), so PPO tends to make answers longer:

In our runs the effect went the other way: replies got shorter, because the EOS rule made cut-off replies the most expensive thing the policy could produce (Section 8.10). Within finished held-out replies, the correlation between our reward model's score and length was +0.24+0.24 for the starting model and −0.10-0.10 after 20 iterations of β=0.05\beta = 0.05 (results/ch8_analysis_stdout.txt). Length bias is a property of the whole set-up, reward model plus length limit plus EOS rule, not of the reward model alone.

Mode collapse. The policy concentrates on a few patterns: the same opening, the same structure, the same phrases, whatever the prompt. Entropy falls, diversity metrics fall, and the KL to the reference rises. Zheng et al. (2023) describe it from 7B runs:

Value-model problems. A value model that predicts badly gives noisy or biased advantages (Section 6.10). Its signatures: explained variance near zero or negative, a value loss that does not fall, and many clipped value updates. A value model initialised from the reward model starts badly on every token except the last (Detail 13 of N+ shows the same picture), so a negative explained variance in the first iterations is expected; one that stays negative is not.

00.020.040.060.0805101520clip fraction00.0050.010.01505101520approx KL (old to new), per token-1-0.500.5105101520PPO iterationvalue model: explained variance0.650.70.750.80.8505101520PPO iterationpolicy entropy, nats per tokenbeta = 0.05no KL penalty3-iteration running mean
PPO's own diagnostics for both runs. Clip fraction and approximate KL are highest in the first iterations, then stay small: each update moves the policy only a little. The value model's explained variance climbs from about -0.7 (a reward model read at every token is a poor critic) to about 0.8. Entropy is noisy and drifts down a little.

The dashboard. Putting it together, these are the numbers to log every iteration, with what our β=0.05\beta = 0.05 run showed for each:

metricwhat it measureshealthyour β=0.05\beta = 0.05 run
reward (proxy)what the policy is paidrises, then flattens−0.02-0.02 to +0.10+0.10 (5-iteration means)
gold or human scorewhat you actually wantrises with the proxy+0.13+0.13 at iteration 10, −0.04-0.04 at 20 (noise ±0.2\pm 0.2)
KL to referencetotal drift from the startgrows slowly, then levels off0 to about 3.5 nats per reply
reply length, share finishedthe cheapest hackstable or explained53 to 31 tokens, 81% to 91% finished
entropydiversity of each next-token choicefalls slowly0.82 to 0.74 nats per token
clip fractionshare of tokens with ratio outside [0.8,1.2][0.8, 1.2]about 0.01 to 0.20.016 on average
approx KL (old to new)size of each updateabout 0.001 to 0.02 per token0.0025 on average
value loss, explained variancequality of the criticloss falls, EV rises above 0EV from −0.70-0.70 to +0.81+0.81
value clip fractionhow often the value update hits the clipsmall0.36: the value model starts far off

The "healthy" ranges are rough rules of thumb from practice, not laws. The first trial of our setup (learning rate 5×10−55 \times 10^{-5}, 4 epochs) had an approx KL of 0.098 and a clip fraction of 0.22 in its first iteration, a sign that each update was too large, and we changed the settings because of it. The value clip fraction of 0.36 is the cost of initialising the value model from a reward model whose per-token outputs are far from the returns (Section 8.6): for many tokens, the value model wanted to move more than 0.2 per update. Its explained variance still rose from negative to 0.81 within 20 iterations.

8.12 What we did not do

To keep the run small enough to repeat in about an hour on one machine, we left out several things that real systems do: an SFT stage of our own (we start from an instruction-tuned model), a learning-rate schedule, batches of hundreds of replies, many seeds, a pretraining loss (PPO-ptx), an adaptive KL controller, and an evaluation by people. Our gold judge is another reward model, which is a stronger and independent reward model but not the truth (Section 7.10 discussed this). Each of these would change the numbers. The mechanics, and the warning signs, are the same.

What comes next

PPO needs a reward model, a value model, a reference model and a sampling loop, and every one of them has to be tuned. The next chapter asks whether all of that is necessary. Direct Preference Optimization (DPO) uses the closed form of the KL-regularised optimum from Section 6.12 to turn the preference data itself into a loss on the policy: no reward model to train, no sampling during training, no value model, and a loss that looks like the reward model loss of Chapter 7. We will derive it, train it on the same UltraFeedback pairs, and compare it with the PPO run of this chapter.

References

Papers

  • Schulman, J., Levine, S., Moritz, P., Jordan, M. I., Abbeel, P. (2015). Trust Region Policy Optimization. ICML 2015. arXiv:1502.05477
  • Schulman, J., Moritz, P., Levine, S., Jordan, M., Abbeel, P. (2015). High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv:1506.02438
  • Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347
  • Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., Irving, G. (2019). Fine-Tuning Language Models from Human Preferences. arXiv:1909.08593
  • Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., Christiano, P. (2020). Learning to summarize from human feedback. NeurIPS 2020. arXiv:2009.01325
  • Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback (InstructGPT). arXiv:2203.02155
  • Bai, Y., Jones, A., Ndousse, K., et al. (2022), Anthropic. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862
  • Huang, S., Noukhovitch, M., Hosseini, A., Rasul, K., Wang, W., Tunstall, L. (2024). The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization. arXiv:2403.17031
  • Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., Madry, A. (2020). Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO. ICLR 2020. arXiv:2005.12729
  • Andrychowicz, M., Raichuk, A., Stańczyk, P., et al. (2021). What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study. ICLR 2021. arXiv:2006.05990
  • Zheng, R., Dou, S., Gao, S., et al. (2023). Secrets of RLHF in Large Language Models Part I: PPO. arXiv:2307.04964
  • Gao, L., Schulman, J., Hilton, J. (2022). Scaling Laws for Reward Model Overoptimization. arXiv:2210.10760
  • Singhal, P., Goyal, T., Xu, J., Durrett, G. (2023). A Long Way to Go: Investigating Length Correlations in RLHF. arXiv:2310.03716
  • Ahmadian, A., Cremer, C., Gallé, M., et al. (2024). Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs. arXiv:2402.14740

Other sources