How Models Are Trained · Part 1 · The Big Picture

Chapter 1 · From next word to assistant

What a language model computes (tokens, next-token probabilities, the chain rule, softmax, cross-entropy, perplexity), the full training pipeline from pretraining to RL with verifiable rewards, and hands-on experiments with Qwen2.5-0.5B base and instruct: the chat template, how few tokens post-training really changes, and a tiny SFT loop.

Goal: by the end of this chapter you can explain what a language model computes, write down the probability of a sentence and the training loss with every symbol named, and work them out with real numbers. You can draw the full training pipeline (pretraining, mid-training, supervised fine-tuning, preference tuning, RL with verifiable rewards) and say, for each stage, what data it uses, what it optimises, how big it is and what it changes. You will have seen a base model and its instruct version answer the same questions, looked at the chat template token by token, measured how many tokens post-training actually changes, and run a small training loop yourself.


1.1 Two models, one question

Here are two small models from the same family. Both have 494 million weights, the same architecture and the same vocabulary. The first, Qwen2.5-0.5B, is a base model: it has only ever been trained to continue text. The second, Qwen2.5-0.5B-Instruct, started as an exact copy of the first and was then trained further to act as an assistant. We ask both the same question, "Who are you?", with greedy decoding (the model always takes its single most likely next token, so there is no randomness and anyone who runs the script gets the same text).

The base model, given the question as plain text, writes:

plain text
 What is your purpose? What is your purpose in life? What is your purpose in life? What is your purpose in life? ...

and keeps going until we cut it off at 200 tokens. The instruct model, given the same question in the format it was trained on, writes:

plain text
I am Qwen, a large language model created by Alibaba Cloud. I am a language model that can generate human-like
text based on the input I receive. ...

and then stops by itself after 87 tokens.

The base model did nothing wrong. It was trained to continue text the way text on the internet continues, and on the internet the line "Who are you?" is often followed by more questions: a list of journaling prompts, a song lyric, a quiz. The instruct model was trained to treat the text as a message from a person and to write one reply, then end its turn.

This chapter is about the distance between those two outputs. What is a language model actually computing? What training turns the first model into the second? How much of the model does that training really change? And how could you do a small piece of that training yourself?

All the experiments in this chapter use these two models because they are small enough to run on a laptop, the base and the instruct version are both published, and the team that made them describes their training in a technical report (Qwen Team, 2024) that we can quote. Every number in the text that does not come from a paper comes from a script in the book's repository, and the script's name is given.

1.2 What a language model is

Before we can say what training changes, we need to be exact about what the thing being trained is. A language model has three parts: a tokenizer that turns text into numbers, a large set of weights that does the computing, and an output that is a probability for every possible next token.

Tokens

A model does not read letters or words. It reads tokens: pieces of text from a fixed list, each with a number (its id). Common words are one token; rare words are split into several pieces; spaces and punctuation are part of tokens too.

Here is the sentence we will use for the next few sections, as the Qwen2.5 tokenizer splits it:

textThe capital of France is Paris.tokensThe785·capital6722·of315·France9625·is374·Paris12095.13ids7 tokens. The model never sees letters, only these ids: rows of a table of 151,936 entries."·" marks a space: most tokens carry their leading space, so " Paris" and "Paris" are different tokens.
The sentence "The capital of France is Paris." becomes 7 tokens. Each token has an id, and the ids are all the model ever sees. Output of ch1_nexttoken.py.

Notice that most tokens start with a space: " capital", " of", " Paris". The token for " Paris" (with a space, id 12095) is a different token from "Paris" (no space). This matters later, when we look at what the model wants to write first in an answer.

Weights

The weights are organised as a transformer: a stack of 24 identical layers, each of which mixes information between positions (attention) and then transforms each position separately (a small feed-forward network). Inside the model every token is a list of 896 numbers (the hidden size). This book does not need the inside of the transformer; for us the model is a function with a very large number of adjustable knobs.

The output: one probability for every possible next token

The model's job is narrow and precise. Given the tokens so far (the context), it outputs a score for every one of the 151,936 entries in its output table, and those scores become probabilities that sum to 1. That list of probabilities is the model's prediction of what comes next.

context (token ids)785 6722 315 9625 374Qwen2.5-0.5B494,032,768 weights24 transformer layershidden size 896one probability per vocabulary entry·Paris Paris: 0.31560.316·______ ______: 0.11460.115·____ ____: 0.06350.064·__ __: 0.05510.055:↵: : 0.05130.051... 151,931 moreTraining = changing the weights so that the right next token gets a higher probability.
A language model as a function. Token ids go in; the 494 million weights compute; out comes one probability for every possible next token. Shown: the five most likely tokens after "The capital of France is", with the real probabilities from Qwen2.5-0.5B.

To generate text, you take that distribution, pick one token (the most likely, for greedy decoding, or a random draw, for sampling), append it to the context, and run the model again. One new token per run. A 200-token answer is 200 runs of the same model, each one seeing one more token than the last. Everything a chat model says, it says this way.

1.3 Next-token prediction, one position at a time

Let us watch the base model read our sentence. At every position it sees the tokens so far and produces a distribution. We can look up the probability it gave to the token that really came next.

1context: "The"actual next·capitalP( capital) = 0.00010.0001model's top-1·followingtop-1 following: 0.6840.684rank of the actual token: 521 · −log P = 9.3162context: "The capital"actual next·ofP( of) = 0.54100.5410model's top-1·oftop-1 of: 0.5410.541rank of the actual token: 1 · −log P = 0.6143context: "The capital of"actual next·FranceP( France) = 0.04160.0416model's top-1·Bertop-1 Ber: 0.1860.186rank of the actual token: 4 · −log P = 3.1794context: "The capital of France"actual next·isP( is) = 0.72110.7211model's top-1·istop-1 is: 0.7210.721rank of the actual token: 1 · −log P = 0.3275context: "The capital of France is"actual next·ParisP( Paris) = 0.31560.3156model's top-1·Paristop-1 Paris: 0.3160.316rank of the actual token: 1 · −log P = 1.1536context: "The capital of France is Paris"actual next.P(.) = 0.46530.4653model's top-1.top-1 .: 0.4650.465rank of the actual token: 1 · −log P = 0.765
Six frames, one per position. Each shows the context, the probability the model gave to the actual next token (upper bar) and the model's own favourite (lower bar). Real numbers from Qwen2.5-0.5B, ch1_nexttoken.py.

Read the frames one at a time; they are a small lesson in what the model has learned.

  1. After "The", the model gives " capital" a probability of only 0.00009. It ranks it 521st. Its favourite is " following" (0.684), because "The following ..." starts a great many documents, lists and exam questions. With one word of context, " capital" is a genuinely unlikely continuation.
  2. After "The capital", " of" is the favourite (0.541). Grammar.
  3. After "The capital of", the model spreads its bets across countries and cities. " France" gets 0.042 and is ranked 4th; the favourite is " Ber" (the first piece of "Berlin", 0.186). This is world knowledge in the form of a distribution: many capitals are plausible here.
  4. After "The capital of France", " is" is the favourite at 0.721.
  5. After "The capital of France is", " Paris" is the favourite, but at only 0.316. This surprises most people. The softmax example below shows where the other 68% went.
  6. After "... is Paris", the full stop is the favourite at 0.465.

No probability is given for the first token, "The". The model needs at least one token of context, and Qwen2.5 does not put a special "beginning of text" token in front of ordinary text, so the first token is simply given.

From scores to probabilities: softmax

The last layer of the model does not output probabilities directly. It outputs one raw score per vocabulary entry, called a logit. A function called softmax turns the scores into probabilities.

P(v∣x<t)=exp⁡(zv)∑u=1Vexp⁡(zu)P(v \mid x_{<t}) = \frac{\exp(z_v)}{\sum_{u=1}^{V} \exp(z_u)}

where:

  • x<tx_{<t} is the context: all tokens before position tt (read "x before t");
  • vv is one candidate token from the vocabulary;
  • zvz_v is the logit the model computed for vv at this position;
  • exp⁡(z)=ez\exp(z) = e^z, with e≈2.718e \approx 2.718, which turns any number into a positive one and turns differences into ratios;
  • VV is the size of the output table, 151,936 here, and the sum runs over every entry uu;
  • P(v∣x<t)P(v \mid x_{<t}) is the probability of vv being the next token, given the context. The vertical bar is read "given".

Dividing by the sum makes all the probabilities add up to exactly 1. Because exp⁡\exp grows so fast, the token with the highest logit gets the largest share, but every token keeps a small positive probability.

Worked example. After "The capital of France is", the five highest logits are:

After "The capital of France is": logits → softmaxlogit ze^(z − max)softmax (all 151,936)softmax (just these 5)·Parislogit 17.84317.84exp 1.00001.000P 0.31560.316P among 5: 0.52580.526·______logit 16.83016.83exp 0.36310.363P 0.11460.115P among 5: 0.19090.191·____logit 16.24016.24exp 0.20130.201P 0.06350.064P among 5: 0.10590.106·__logit 16.09816.10exp 0.17470.175P 0.05510.055P among 5: 0.09180.092:↵logit 16.02716.03exp 0.16260.163P 0.05130.051P among 5: 0.08550.086Logit bars start at 15.5 to show the gaps. A logit 1.0 higher means e ≈ 2.72 times more probable.These five take 60% of the probability; the other 151,931 tokens share the remaining 40%.
Softmax by hand on the top five candidates. Left to right: the logits, the exponentials after subtracting the largest logit, the true softmax probabilities over all 151,936 entries, and what softmax would give if these five were the only options. ch1_nexttoken.py.

Take " Paris" (logit 17.843) and " ______" (logit 16.830, a fill-in-the-blank line). The difference is 1.013, so " Paris" should be e1.013=2.75e^{1.013} = 2.75 times more probable. The true probabilities are 0.3156 and 0.1146, and 0.3156 / 0.1146 = 2.75. Softmax over all entries does exactly what the formula says.

If the vocabulary contained only these five tokens, the denominator would be 1+0.363+0.201+0.175+0.163=1.9021 + 0.363 + 0.201 + 0.175 + 0.163 = 1.902 (after subtracting the largest logit, which does not change the result), and " Paris" would get 1/1.902=0.5261 / 1.902 = 0.526. In the real model it gets 0.316, because the other 151,931 entries together hold 40% of the probability. Each of them is tiny, but there are a lot of them.

Now look at what the model's top five are: " Paris", then three lengths of blank line (" ______", " ____", " __"), then a colon and a line break. A base model is a mirror of its training data, and on the web the words "The capital of France is" appear very often in quizzes and worksheets, followed by a blank to fill in. Together the blanks get 23% of the probability. The model knows the answer is Paris; it is also, correctly, unsure whether this document wants the answer or the question. Keep this in mind: it is a big part of what post-training will change.

Here is the code that produced these numbers, simplified from ch1_nexttoken.py:

python
tok = AutoTokenizer.from_pretrained('Qwen/Qwen2.5-0.5B')
model = AutoModelForCausalLM.from_pretrained('Qwen/Qwen2.5-0.5B').to('mps').eval()

ids = tok('The capital of France is Paris.')['input_ids']        # [785, 6722, 315, 9625, 374, 12095, 13]
with torch.no_grad():
    logits = model(torch.tensor([ids], device='mps')).logits[0]  # shape [7, 151936]: one row of scores per position
probs = torch.softmax(logits, dim=-1)                            # each row now sums to 1

for t in range(len(ids) - 1):
    p = probs[t, ids[t + 1]]          # row t predicts the token at position t+1
    print(tok.decode(ids[:t + 1]), '->', tok.decode([ids[t + 1]]), float(p))

Line by line: the tokenizer turns the sentence into 7 ids. One forward pass of the model (model(...)) computes logits for all 7 positions at once: row 0 is the prediction after "The", row 1 after "The capital", and so on. This is the trick that makes training efficient: one pass over a sequence gives a prediction, and therefore a training signal, at every position. torch.softmax(..., dim=-1) applies the softmax formula to each row. Finally, row tt is the prediction for position t+1t+1, so we look up the probability of the token that actually sits there. The torch.no_grad() block tells PyTorch we are only reading, not training, so it does not keep the extra bookkeeping needed for gradients.

1.4 The probability of a whole sentence: the chain rule

A language model gives probabilities for one next token. But we often want the probability of a whole sentence: to compare two answers, to score how well the model fits some text, or to train it. Probability theory gives a rule for exactly this.

For a sequence of tokens x1,x2,…,xTx_1, x_2, \dots, x_T:

P(x1,…,xT)=P(x1)∏t=2TP(xt∣x1,…,xt−1)P(x_1, \dots, x_T) = P(x_1) \prod_{t=2}^{T} P(x_t \mid x_1, \dots, x_{t-1})

where:

  • xtx_t is the token at position tt, and TT is the number of tokens;
  • P(x1)P(x_1) is the probability of the first token on its own;
  • ∏\prod (capital Greek "pi") means "multiply together", here for every position from 2 to TT;
  • P(xt∣x1,…,xt−1)P(x_t \mid x_1, \dots, x_{t-1}) is the model's next-token probability at position tt, exactly the number in the frames above.

A language model is called autoregressive because it uses this rule: it only ever predicts one token given the earlier ones, and the probability of anything longer is the product of those predictions.

Worked example. Our sentence has 7 tokens. Since the model does not score the first one, we compute the probability of the other six given "The":

P(capital of France is Paris . | The) = product of six next-token probabilities·capital0.0001log = -9.316×·of0.5410log = -0.614×·France0.0416log = -3.179×·is0.7211log = -0.327×·Paris0.3156log = -1.153×.0.4653log = -0.765product = 2.145e-07sum of logs = -15.355loss = 15.355 / 6 = 2.559Tiny numbers multiply into tinier ones; adding logs gives the same information without underflow.The single unlikely token " capital" (P = 0.0001) contributes 9.32 of the 15.35 total.
The chain rule on real numbers: six next-token probabilities multiply to 2.1 × 10⁻⁷. The logs add up to −15.355, which divided by 6 gives the average loss of 2.559. ch1_nexttoken.py.
0.00009×0.541×0.0416×0.721×0.316×0.465=2.1×10−70.00009 \times 0.541 \times 0.0416 \times 0.721 \times 0.316 \times 0.465 = 2.1 \times 10^{-7}

About two in ten million. That sounds absurdly small for a true and ordinary sentence, but it is the right scale: there are billions of possible six-token continuations of "The", and this is one of them. Probabilities of whole sentences are always tiny, which is why nobody works with them directly. They work with logarithms.

Why logs

The log of a product is the sum of the logs: log⁡(a×b)=log⁡a+log⁡b\log(a \times b) = \log a + \log b. So:

log⁡P(x2,…,xT∣x1)=∑t=2Tlog⁡P(xt∣x<t)\log P(x_2, \dots, x_T \mid x_1) = \sum_{t=2}^{T} \log P(x_t \mid x_{<t})

where log⁡\log is the natural logarithm (base ee), and every term is negative or zero because every probability is at most 1.

For our sentence the six logs are −9.316, −0.614, −3.179, −0.327, −1.153 and −0.765, and their sum is −15.355. Taking e−15.355e^{-15.355} gives back 2.1×10−72.1 \times 10^{-7}. Sums are easier to work with than products, and computers cannot store numbers like 10−500010^{-5000} (the probability of a long document) without special tricks, while a sum like −11,500 is no problem.

1.5 Cross-entropy loss and perplexity

Training needs one number that says how badly the model did, so that it can change the weights to make that number smaller. For language models the number is the average negative log-probability of the actual tokens.

L(θ)=−1T−1∑t=2Tlog⁡Pθ(xt∣x<t)\mathcal{L}(\theta) = -\frac{1}{T-1} \sum_{t=2}^{T} \log P_\theta(x_t \mid x_{<t})

where:

  • θ\theta (Greek "theta") stands for all the weights of the model, and PθP_\theta is the model's probability with those weights;
  • L\mathcal{L} (a curly L) is the loss: the number training tries to make small;
  • the minus sign turns the negative logs into a positive number;
  • dividing by T−1T-1 (the number of predictions) makes it an average per token, so long and short texts are comparable.

The curve below is −log⁡P-\log P for every possible PP, with our six real tokens marked on it.

024681000.250.50.751 capital: P=0.0001, -log P=9.316·capital (9.32) of: P=0.5410, -log P=0.614·of (0.61) France: P=0.0416, -log P=3.179·France (3.18) is: P=0.7211, -log P=0.327·is (0.33) Paris: P=0.3156, -log P=1.153·Paris (1.15).: P=0.4653, -log P=0.765. (0.77)probability the model gave to the actual next token, Ploss = −log P (nats)
The per-token loss −log P. The six tokens of our sentence sit on the curve. The loss is flat and near zero when the model is confident and right, and rises steeply as P falls. One token, " capital", supplies most of the total. ch1_nexttoken.py.

Worked example. The six losses are 9.316, 0.614, 3.179, 0.327, 1.153 and 0.765. Their sum is 15.355 and their average is 15.355 / 6 = 2.559 nats per token. The library computes the same number in one call; the script checks this:

python
out = model(torch.tensor([ids]), labels=torch.tensor([ids]))
print(out.loss)      # 2.559, identical to our hand computation

When you pass labels, the library shifts them by one position for you (the label for row tt is the token at t+1t+1), computes the softmax and the negative log at every position, and averages. That one line is the pretraining objective. Everything in Section 1.10 is built on it.

Notice how unequal the six terms are. " capital" alone is 9.3 of the 15.4. A single surprising token dominates the loss of a short sentence. Over billions of tokens these surprises average out, and the loss becomes a smooth measure of how well the model predicts text.

Perplexity

The loss in nats is hard to picture. Perplexity turns it back into something countable.

PPL=exp⁡(L)\text{PPL} = \exp(\mathcal{L})

where L\mathcal{L} is the average cross-entropy loss in nats, and exp⁡\exp undoes the log.

For our sentence, e2.559=12.9e^{2.559} = 12.9: on average, the model was as unsure as if it were picking among about 13 equally likely tokens at each step. The script also scores two more sentences:

SentenceLoss (nats/token)Perplexity
The capital of France is Paris.2.55912.9
The cat sat on the mat.2.95719.2
Purple ideas sleep furiously under quiet spoons.7.6682,139

The first two are ordinary English and get low perplexity. The third is grammatical nonsense, and the model is very surprised by almost every word. Perplexity on held-out text is the main number people watch during pretraining, and Chapter 3 discusses its limits (in short: it tells you how well a model predicts text, not how good its answers are).

1.6 The training pipeline: from base model to assistant

Everything above describes a single model and a single loss. Real models are trained in stages. Each stage starts from the weights the previous stage produced, uses different data, and often a different objective. Here is the whole pipeline as it looks in the technical reports of 2024 and 2025 (Qwen2.5, Llama 3, OLMo 2, Tulu 3, DeepSeek-R1). The numbers in the "scale" boxes are quoted from those reports.

One model, five training stages (each starts from the weights the previous stage left)Pretrainingdataweb pages, books,code, papersobjectivepredict the nexttoken, everywherescale (published)Qwen2.5: 18T tokensLlama 3: 15.6T tokenswhat it changesknowledge, grammar,skillsMid-trainingdatacurated high-qualitymaths, code, long textobjectivesame objective,learning rate to 0scale (published)OLMo 2 7B: 3 runs ×50B (3.7% of 4.05T)what it changesfills gaps, longercontextSFTdataprompt + one goodanswer, chat formatobjectivenext token, on theanswer tokens onlyscale (published)Qwen2.5: 1M+ examplesInstructGPT: 13k promptswhat it changesformat, turn-taking,when to stopPreference tuningdataprompt + preferred+ rejected answerobjectiveRLHF (reward model+ PPO), or DPOscale (published)Qwen2.5 DPO: ~150k pairsLlama 3: 6 roundswhat it changeswhich answer peoplelike moreRL, verifiable rewarddataproblems with acheckable answerobjectivereward 1 if the finalanswer is rightscale (published)DeepSeek-R1-Zero: AIME15.6% → 77.9%what it changeslonger, more carefulreasoningpre-training → the "base" model (Qwen2.5-0.5B)post-training → the "instruct" model (Qwen2.5-0.5B-Instruct)compute: InstructGPT SFT 4.9 and PPO 60 petaflop/s-days, against 3,640 for pretraining GPT-3
The training pipeline. Pretraining and mid-training produce a base model; supervised fine-tuning, preference tuning and RL with verifiable rewards turn it into an assistant. For each stage: its data, its objective, its published scale, and what it changes. Sources: Qwen Team (2024), Grattafiori et al. (2024), Team OLMo (2024), Ouyang et al. (2022), DeepSeek-AI (2025).

The Llama 3 report describes the two halves in two short paragraphs that are worth reading in the original:

The rest of this section walks through the five stages. For each one: the data, the objective (as an equation preview; later parts of the book derive each one properly), the scale, and what it changes. Part 1 of this book (this chapter and the next two) is the big picture. Later parts take supervised fine-tuning, RLHF, DPO and RL for reasoning one at a time, in depth, with code.

Stage 1: pretraining

Data. As much good text as can be found: web pages (filtered and de-duplicated), books, scientific papers, code, maths, in many languages. The Llama 3 report gives its final mix as roughly 50% general knowledge, 25% maths and reasoning, 17% code and 8% multilingual tokens (Grattafiori et al., 2024, Section 3.1.2).

Objective. Exactly the loss of Section 1.5, averaged over every position of every document in the training set:

Lpretrain(θ)=− Ex∼D[1T∑t=1Tlog⁡Pθ(xt∣x<t)]\mathcal{L}_{\text{pretrain}}(\theta) = -\,\mathbb{E}_{x \sim \mathcal{D}} \left[ \frac{1}{T} \sum_{t=1}^{T} \log P_\theta(x_t \mid x_{<t}) \right]

where:

  • D\mathcal{D} is the training corpus, and x∼Dx \sim \mathcal{D} means "a document xx drawn from it";
  • E\mathbb{E} means "the average over" (the expectation), here over documents;
  • the inside of the brackets is the per-token average log-probability of one document, as before.

Nothing is labelled by a person. The text is its own label: every token is the answer to the question "what comes next?" for the tokens before it. This is called self-supervised learning, and it is why pretraining can use trillions of tokens.

Scale. This stage is where nearly all the compute goes.

ModelParametersPretraining tokensSource
GPT-3 (2020)175B300BBrown et al. (2020), Table 2.1
Llama 3 (2024)405B15.6TGrattafiori et al. (2024), Section 1
Qwen2.5 (2024)0.5B to 72B18TQwen Team (2024), Section 3
OLMo 2 7B (2024)7B3.9T (plus mid-training)Team OLMo (2024), Section 2

The Llama 3 flagship used 3.8×10253.8 \times 10^{25} floating-point operations for pretraining, "almost 50× more than the largest version of Llama 2" (Grattafiori et al., 2024, Section 1).

What it changes. Everything, from random numbers to a model that knows grammar, facts, styles, code, some arithmetic and a great deal of the structure of the world as described in text. The output is a base model. It can already do many tasks if you phrase them as text to be continued, as GPT-3 showed:

Stage 2: mid-training

A newer name for an old idea. Near the end of pretraining, the data mix is changed: the share of high-quality text, maths, code and long documents goes up, and the learning rate is lowered towards zero (annealing). The objective stays the same next-token loss.

Scale. OLMo 2 7B is pretrained on 3.90 trillion tokens, then mid-trained three separate times on 50 billion tokens each, and the three results are averaged: 4.05 trillion tokens in total (OLMo et al., 2024, Section 2.3). So mid-training is about 3.7% of the tokens. Llama 3 anneals on its final 40 million tokens, and reports that annealing on high-quality maths and code improved the 8B model on the GSM8k and MATH benchmarks by 24.0% and 6.4% (Grattafiori et al., 2024, Section 3.1.3).

What it changes. Sharper maths and code, longer context, and, often, some familiarity with question-and-answer formats. This matters for our experiments: the Qwen2.5 report says its pretraining data includes synthetic data "particularly in mathematics, code, and knowledge domains" generated with Qwen2-72B-Instruct and Qwen2-Math-72B-Instruct (Qwen Team, 2024, Section 3.1). So the Qwen2.5 base model has already read a lot of text written by an assistant. You will see the effect in Section 1.7.

Stage 3: supervised fine-tuning (SFT)

Data. Pairs of a prompt (an instruction or question, possibly with earlier turns of a conversation) and a response written the way the final assistant should write. The responses are written by people, generated by a stronger model and filtered, or both. Every example is put into the chat template (Section 1.8) so the model learns where a turn begins and ends.

Objective. The same next-token loss, with one change: only the tokens of the response count. The prompt is in the context but is not trained on.

LSFT(θ)=− E(x,y)∼DSFT[1∣y∣∑t=1∣y∣log⁡Pθ(yt∣x,y<t)]\mathcal{L}_{\text{SFT}}(\theta) = -\,\mathbb{E}_{(x, y) \sim \mathcal{D}_{\text{SFT}}} \left[ \frac{1}{|y|} \sum_{t=1}^{|y|} \log P_\theta(y_t \mid x, y_{<t}) \right]

where:

  • xx is the prompt (with the chat template around it) and yy is the response, y1…y∣y∣y_1 \dots y_{|y|};
  • ∣y∣|y| is the number of tokens in the response;
  • Pθ(yt∣x,y<t)P_\theta(y_t \mid x, y_{<t}) is the probability of the tt-th response token given the whole prompt and the response so far;
  • DSFT\mathcal{D}_{\text{SFT}} is the set of prompt-response pairs.

Compare it with the pretraining loss: it is the same formula with the sum restricted to the response. In code, this restriction is a single line that sets the label of every prompt token to −100 (Section 1.10).

Scale. Tiny compared with pretraining. InstructGPT's SFT set had about 13,000 training prompts (Ouyang et al., 2022, Section 3.2). Qwen2.5's has "over 1 million" examples, trained for two epochs with sequences up to 32,768 tokens (Qwen Team, 2024, Section 4.1). LIMA trained a 65B model on just 1,000 carefully chosen examples (Zhou et al., 2023). Even a million examples of a thousand tokens each is about a billion tokens: less than 0.01% of an 18-trillion-token pretraining run.

What it changes. Format and behaviour: answering instead of continuing, the turn structure, the tone, when to stop, when to refuse, how long to be. Whether it also teaches new knowledge is exactly the question of Section 1.9.

Stage 4: preference tuning (RLHF, DPO)

SFT can only show the model good answers. It cannot tell the model which of two decent answers is better, or punish a bad habit directly. Preference tuning does this.

Data. For each prompt, two or more responses (usually sampled from the model itself) and a judgement of which one is better, made by a person or by another model. A preference pair is (prompt, chosen response, rejected response).

Objective, version 1: RLHF. First train a reward model rϕ(x,y)r_\phi(x, y) that gives a score to a response, using the preference pairs. Then improve the language model (now called the policy πθ\pi_\theta) with reinforcement learning to get high reward, while staying close to the SFT model:

max⁡θ  Ex∼D, y∼πθ(⋅∣x)[rϕ(x,y)]  −  β KL(πθ ∥ πSFT)\max_{\theta} \; \mathbb{E}_{x \sim \mathcal{D},\, y \sim \pi_\theta(\cdot \mid x)} \Big[ r_\phi(x, y) \Big] \;-\; \beta \, \mathrm{KL}\big(\pi_\theta \,\|\, \pi_{\text{SFT}}\big)

where:

  • πθ(y∣x)\pi_\theta(y \mid x) is the language model being trained, seen as a policy: a rule for choosing outputs; y∼πθ(⋅∣x)y \sim \pi_\theta(\cdot \mid x) means a response sampled from it;
  • rϕ(x,y)r_\phi(x, y) is the reward model's score, with its own weights ϕ\phi (Greek "phi");
  • πSFT\pi_{\text{SFT}} is the frozen SFT model, the starting point;
  • KL(⋅ ∥ ⋅)\mathrm{KL}(\cdot \,\|\, \cdot) is the KL divergence, a measure of how different two distributions are (we define it properly and compute it in Section 1.9);
  • β\beta (Greek "beta") is a number that sets how strongly the model is held near the SFT model.

The KL term is the leash: without it the model finds strange outputs that fool the reward model. InstructGPT is the paper that made this recipe famous:

Objective, version 2: DPO. Rafailov et al. (2023) showed that the same goal can be reached without a separate reward model and without reinforcement learning, with a loss computed directly on preference pairs:

LDPO(θ)=− E(x,yw,yl)[log⁡σ ⁣(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_{\text{DPO}}(\theta) = -\,\mathbb{E}_{(x, y_w, y_l)} \left[ \log \sigma\!\left( \beta \log \frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \beta \log \frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)} \right) \right]

where:

  • ywy_w is the chosen ("winning") response and yly_l the rejected ("losing") one;
  • πref\pi_{\text{ref}} is a frozen reference model, usually the SFT model;
  • each fraction compares how likely the trained model makes a response with how likely the reference model makes it;
  • σ\sigma is the sigmoid function, σ(z)=1/(1+e−z)\sigma(z) = 1 / (1 + e^{-z}), which squashes any number into the range 0 to 1;
  • β\beta again controls how far the model may move from the reference.

In words: raise the probability of the chosen answer and lower the probability of the rejected one, both measured relative to where the model started. No sampling, no reward model; just a loss on log-probabilities, like SFT.

Scale. InstructGPT's reward model used 33,000 training prompts and its PPO stage 31,000 prompts (Ouyang et al., 2022, Section 3.2). Qwen2.5's offline stage used about 150,000 DPO training pairs for one epoch (Qwen Team, 2024, Section 4.2), followed by online RL with GRPO against a reward model (Section 4.3). Llama 3 ran SFT and DPO in six rounds, each round collecting new preference data with the best model so far:

What it changes. Which of the answers the model could already give it actually prefers: more helpful, better organised, more careful, less harmful. InstructGPT's labelers preferred outputs of the 175B InstructGPT over the 175B GPT-3 85 ± 3% of the time, and preferred the 1.3B InstructGPT over the 175B GPT-3, "despite having over 100x fewer parameters" (Ouyang et al., 2022, Section 1).

Stage 5: RL with verifiable rewards (RLVR)

Data. Problems whose answer a program can check: maths with a known final number, code with unit tests, puzzles with a unique solution, instructions with checkable constraints ("answer in exactly three sentences").

Objective. The same reinforcement-learning idea as RLHF, but the reward is not a learned model of human taste. It is a rule:

r(x,y)={1if the final answer in y is correct for problem x0otherwiser(x, y) = \begin{cases} 1 & \text{if the final answer in } y \text{ is correct for problem } x \\ 0 & \text{otherwise} \end{cases}

where xx is the problem, yy is the model's full response (reasoning and answer), and "correct" is decided by a program: compare the boxed number with the reference, run the tests.

What it changes. Longer and more careful step-by-step reasoning, self-checking, and accuracy on exactly the kinds of problems that can be checked. Qwen2.5 uses the same kind of automatic checking ("execution feedback and answer matching") to build the preference pairs for maths and code in its offline RL stage (Qwen Team, 2024, Section 4.2). A later part of the book is devoted to RL for reasoning.

The whole picture in one table

StageDataLoss or rewardTypical sizeWhat it changesCovered in depth
Pretrainingtrillions of tokens of textnext-token, all tokens15T to 18T tokensknowledge, language, skillsSections 1.2 to 1.5
Mid-trainingcurated high-quality and synthetic textnext-token, all tokenstens to hundreds of billions of tokensmaths, code, long context, Q&A familiaritythis section
SFTprompt + ideal responsenext-token, response only1,000 to a few million examplesformat, turn-taking, stopping, tonea later part on SFT
Preference tuningprompt + chosen + rejectedreward model + RL (RLHF), or DPO losstens to hundreds of thousands of pairswhich answer it preferslater parts on RLHF and DPO
RLVRproblems with checkable answers1 if correct, else 0thousands to hundreds of thousands of problemsreasoning, accuracy on checkable tasksa later part on RL for reasoning

What post-training costs

If you remember one fact about the pipeline, make it this one. Post-training is a small fraction of the cost of pretraining, and yet it is what turns a text continuer into an assistant.

Training compute for the 175B models, petaflop/s-days (log scale)1101001,000GPT-3 pretrainingGPT-3 pretraining: 3,640.03,640InstructGPT PPO-ptxInstructGPT PPO-ptx: 60.060InstructGPT SFTInstructGPT SFT: 4.94.9Source: Ouyang et al. (2022), Section 5.1. SFT used about 0.13% of the pretraining compute; RLHF about 1.6%.
Compute for the 175B models in petaflop/s-days, on a log scale, from Ouyang et al. (2022), Section 5.1. Each gridline is ten times the previous one.

That imbalance raises the question at the heart of this chapter. If post-training uses so little compute and so little data, what can it possibly be changing? We will answer it by measurement. But first, let us see the change with our own eyes.

1.7 Hands-on: the base model and the instruct model, side by side

The script ch1_base_vs_instruct.py gives five prompts to the two models in three settings:

  1. base, raw text: the base model gets just the question, as plain text, the way you would test a pretrained model;
  2. base, chat template: the base model gets the question wrapped in the chat format that the instruct model was trained on (Section 1.8 shows it token by token);
  3. instruct, chat template: the instruct model gets the chat format, as its makers intended.

Decoding is greedy, at most 200 new tokens, and both models stop if they produce either of the two end tokens, <|endoftext|> or <|im_end|>. The core of the script, simplified:

python
def greedy(tok, model, ids, max_new_tokens=200):
    stop_ids = [tok.convert_tokens_to_ids(t) for t in ['<|endoftext|>', '<|im_end|>']]  # same stop rule for both
    out = model.generate(torch.tensor([ids]), max_new_tokens=max_new_tokens, do_sample=False,
                         eos_token_id=stop_ids)
    return out[0, len(ids):].tolist()                     # only the new tokens

raw  = tok_b(q)['input_ids']                               # setting 1: the question as plain text
chat = tok_i.apply_chat_template([{'role': 'user', 'content': q}],
                                 add_generation_prompt=True, tokenize=True)   # settings 2 and 3

do_sample=False is greedy decoding: no randomness, so the outputs below are exactly reproducible. apply_chat_template builds the chat format from a list of messages; add_generation_prompt=True adds the line that opens the assistant's turn, so the model's next token is the first token of its answer. The two models share one tokenizer (the script checks this), so the same text becomes the same ids for both.

prompt: "What is the capital of France?"base, raw text8 tokens, stopped by itselfThe capital of France is Paris.[END]base, chat template200 tokens, cut at 200The capital of France is Paris.navigationOptions ↵navigationOptions ↵ You are ahelpful assistant.icode ↵ icode ↵What is the capital of France?icode↵ icode ↵ You are a helpful...instruct, chat template8 tokens, stopped by itselfThe capital of France is Paris.[END]prompt: "Who are you?"base, raw text200 tokens, cut at 200What is your purpose? What is yourpurpose in life? What is yourpurpose in life? What is yourpurpose in life? What is yourpurpose in life? What is yourpurpose in life? What is your...base, chat template200 tokens, cut at 200I am a system that can help you withvarious tasks.🤨 ↵ ughtyuser ↵ Whatcan you do for me?ughtyassistant ↵ughtyuser ↵ I can help you withtasks like organizing your files,creating a to-do list, or even just...instruct, chat template87 tokens, stopped by itselfI am Qwen, a large language modelcreated by Alibaba Cloud. I am alanguage model that can generatehuman-like text based on the input Ireceive. I am designed to assistwith a wide range of tasks, from...
Two of the five prompts, in all three settings. Red badges mark outputs that never ended and were cut at 200 tokens. Greedy decoding, ch1_base_vs_instruct.py.

Here is everything the script found, prompt by prompt.

"What is the capital of France?" The base model, given the raw question, answers " The capital of France is Paris." and stops. The instruct model gives the same sentence and stops. Inside the chat template, the base model also starts with "The capital of France is Paris." but then cannot end its turn: it produces odd tokens (" navigationOptions", "icode") and starts repeating "You are a helpful assistant." and the question, over and over, until it is cut off.

"Write a haiku about rain." Raw, the base model writes three short lines and stops: "Raindrops fall, / Soft and gentle, / Nature's gift to us all." The instruct model writes "Raindrops dance on the ground, / A symphony of nature's music, / Peace and tranquility." and stops. (Neither has the 5-7-5 syllable pattern of a real haiku; a model of half a billion weights does not count syllables well.) In the chat template the base model writes two lines and then falls into a loop of "reeze" and the prompt.

"A shop sells pencils at 3 for 45 cents. How much do 7 pencils cost?" All three settings work it out step by step: 45 / 3 = 15 cents per pencil, 7 × 15 = 105 cents = 1.05 dollars. Both base settings are still talking at 200 tokens. The instruct model writes its working in LaTeX and finishes after 197 tokens with the answer in a box, \boxed{1.05}, then ends its turn. The box is a habit from post-training on maths: as Section 1.6 showed, a boxed final answer is what an automatic checker looks for.

"Give me three tips for sleeping better." All three give three sensible tips. The instruct model opens with "Certainly! Here are three tips for improving your sleep quality:", gives each tip a bold heading, adds a closing summary, and is still going at 200 tokens; the base model in the chat template opens with "Sure, here are three tips for sleeping better:" and stops by itself; the raw base model just starts with "1." and also stops.

"Who are you?" The raw base model repeats "What is your purpose in life?" until cut off. In the chat template, it says "I am a system that can help you with various tasks." and then invents a whole conversation, writing the user's lines as well as its own. The instruct model says "I am Qwen, a large language model created by Alibaba Cloud..." and stops after 87 tokens.

What this shows, honestly

This is not the dramatic "base model is useless, instruct model is brilliant" picture you may have expected, and it is worth understanding why.

  1. The base model already knows the answers. Paris, 105 cents, sensible sleep tips: the knowledge is all there before any post-training. This is the first piece of evidence for the claim we will test properly in Section 1.9.
  2. Qwen2.5-0.5B is not a "pure" base model. On three of the five prompts, the raw base model answered the question instead of continuing the text in some other way. In the chat template, it even opened with "Sure, here are three tips". As Section 1.6 noted, the Qwen2.5 pretraining data includes synthetic data written by Qwen2 instruct models. Older base models, such as GPT-3 in 2020, were much less willing to answer a bare question. Today's line between "base" and "instruct" is blurrier than the names suggest.
  3. The clearest difference is turn-taking: knowing when to stop. The instruct model ended its turn by itself on four of five prompts (the fifth was the long, formatted sleep-tips answer, which hit the limit). The base model, in the chat format, ended only the sleep-tips answer, and only with <|endoftext|>, never with <|im_end|>. Everywhere else it rambled, looped, or wrote the user's next message for them.
  4. Identity comes from post-training. "I am Qwen, created by Alibaba Cloud" is not something the base model says. It comes from the default system prompt ("You are Qwen, created by Alibaba Cloud. You are a helpful assistant.") together with the post-training that taught the model to take that message seriously.
  5. Post-training did not fix everything. Neither model writes a correct haiku. It is a small model.

Point 3 deserves a closer look, because the reason is visible in the tokens themselves.

1.8 The chat template, token by token

A chat model does not see "messages". It sees one long sequence of tokens, exactly like a base model does. The conversation structure is written into that sequence with a few special tokens. Here is our first prompt exactly as the instruct model reads it:

"What is the capital of France?" as the instruct model reads it: 36 tokens<|im_start|>system↵You·are·Qwen,·created·by·Alibaba·Cloud.·You·are·a·helpful·assistant.<|im_end|>↵<|im_start|>user↵What·is·the·capital·of·France?<|im_end|>↵<|im_start|>assistant↵special token: one id, never produced by ordinary text (<|im_start|> = 151644, <|im_end|> = 151645)role name and line breakordinary text (the default system prompt, then the question)The template ends with "<|im_start|>assistant↵": the model's job is to write the next turn and close it with <|im_end|>.
The chat template for "What is the capital of France?", all 36 tokens. Highlighted blocks are special tokens; each marks the start or end of a turn. ch1_base_vs_instruct.py.

As text, the same thing looks like this:

plain text
<|im_start|>system
You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>
<|im_start|>user
What is the capital of France?<|im_end|>
<|im_start|>assistant

Reading the 36 tokens in order:

  • <|im_start|> system ↵ opens the system turn. The role name is ordinary text ("system" is token 8948, a normal word).
  • 16 tokens of the default system prompt follow, then <|im_end|> ↵ closes the turn.
  • <|im_start|> user ↵, the 7 tokens of the question, <|im_end|> ↵: the user turn.
  • <|im_start|> assistant ↵ opens the assistant turn, and the sequence ends there. The next token the model writes is the first token of its answer.

When the answer is finished, a well-trained model writes <|im_end|>. The software that runs the model watches for that token, stops generating, and shows you the answer. If the model never writes it, the software keeps asking for more tokens, and the model keeps going: it closes the turn with something else, opens a new "user" turn, and writes a conversation with itself. That is exactly what the base model did in Section 1.7.

Why the template matters

The base model never learned to end a turn. The <|im_end|> token exists in the base model's vocabulary, but the base model gives it essentially zero probability where the instruct model writes it. Section 1.9 measures this: across 24 answers that the instruct model ended with <|im_end|>, the instruct model gave that token an average probability of 0.81, and the base model, reading exactly the same context, gave it at most 0.00003. Where the base model wanted to stop, it put its probability on <|endoftext|>, the end-of-document token it saw at the end of every pretraining document. Teaching the model this one token, and the turn structure around it, is a large part of what SFT does.

Use the template the model was trained with. If you send an instruct model text without its template, or with another family's template, it is reading a format it rarely saw in training, and quality drops in ways that are hard to spot. Libraries ship the template with the tokenizer (apply_chat_template) precisely so you do not have to type it by hand. Getting it subtly wrong (a missing line break, a missing add_generation_prompt) is one of the most common bugs when people fine-tune or serve models.

The template is part of the training data format. In SFT (Section 1.10) every training example is written in this template, and the loss is computed only on the tokens of the assistant's turn, including its closing <|im_end|>. That is how the model learns both what to say and when to stop.

1.9 How much does post-training really change?

Section 1.6 ended with a puzzle: post-training uses a tiny fraction of the compute and data of pretraining, yet it turns a text continuer into an assistant. Section 1.7 added a clue: the base model already knew the answers; what it lacked was the manners. Two papers from 2023 turned this clue into a hypothesis and a measurement.

The superficial alignment hypothesis

LIMA's evidence was indirect: a model fine-tuned on 1,000 examples, with no preference tuning at all, produced answers that human raters judged equivalent to or better than GPT-4's in 43% of cases (Zhou et al., 2023, Section 1). That shows a little SFT goes a long way. It does not show what changed inside the model.

Measuring the change token by token

Lin et al. (2023) found a direct way to look. Take the instruct model's answer. At every position, ask the base model what it would have written next, given exactly the same context. If post-training changed everything, the base model would disagree often. If it only changed the surface, the base model would mostly agree.

Measuring the shift at one position tcontext = prompt + o₁ … o₍ₜ₋₁₎(the instruct model's own answer)instruct: greedy oₜbase: rank of oₜunshifted: rank 1: both models agreemarginal: rank 2 or 3shifted: rank > 3Also at every position: KL(P_instruct ‖ P_base), how different the two whole distributions are.
The measurement at one position t. The instruct model's greedy token oₜ is compared with the base model's ranking for the same context, and the position is classed as unshifted, marginal or shifted. The KL divergence between the two full distributions is recorded too.

In symbols, at position tt of the instruct model's answer o=o1…oTo = o_1 \dots o_T to a query qq:

ηt=rank of ot in Pbase(⋅∣q,o<t),classt={unshiftedηt=1marginalηt∈{2,3}shiftedηt>3\eta_t = \text{rank of } o_t \text{ in } P_{\text{base}}(\cdot \mid q, o_{<t}), \qquad \text{class}_t = \begin{cases} \text{unshifted} & \eta_t = 1 \\ \text{marginal} & \eta_t \in \{2, 3\} \\ \text{shifted} & \eta_t > 3 \end{cases}

where:

  • qq is the user's query in the chat template, and o<to_{<t} is the instruct model's answer up to position t−1t-1;
  • oto_t is the token the instruct model chose at position tt (with greedy decoding, its own top-1);
  • Pbase(⋅∣q,o<t)P_{\text{base}}(\cdot \mid q, o_{<t}) is the base model's next-token distribution for the same context, the dot meaning "over all tokens";
  • ηt\eta_t (Greek "eta") is the base rank: 1 if the base model would also have chosen oto_t, 2 if it was its second choice, and so on.

Our experiment

The script ch1_token_shift.py runs this on Qwen2.5-0.5B and Qwen2.5-0.5B-Instruct, with 40 prompts of six kinds that we wrote for this chapter: knowledge questions ("Why is the sky blue?"), how-to and advice ("How do I make a cup of tea?"), creative writing ("Write a short poem about the ocean."), maths ("Is 97 a prime number?"), code ("Write a Python function that reverses a string.") and conversation ("Hello! How are you today?", "Tell me a joke."). The instruct model answers each with greedy decoding (up to 200 tokens). That gives 5,223 answer tokens.

python
for q in PROMPTS:
    ctx = chat(q)                                   # the chat-template tokens of the query
    ans, stopped = greedy(tok, inst, ctx, 200)      # the instruct model's answer o_1 ... o_T
    full = ctx + ans
    li = log_softmax(inst(full).logits)[len(ctx)-1 : len(full)-1]   # row t predicts ans[t]
    lb = log_softmax(base(full).logits)[len(ctx)-1 : len(full)-1]   # same rows, from the base model
    for t, a in enumerate(ans):
        rank = (lb[t] > lb[t, a]).sum() + 1                          # 1 = base model's top-1
        kl = (li[t].exp() * (li[t] - lb[t])).sum()                   # KL(P_instruct || P_base), Section below

(simplified from ch1_token_shift.py). The trick is the same as in Section 1.3: a single forward pass over prompt plus answer gives the distribution at every position at once, so each model runs once per prompt, not once per token. The slice [len(ctx)-1 : len(full)-1] picks the rows that predict the answer tokens. The rank is computed by counting how many tokens the base model scored higher than the instruct model's choice.

Qwen2.5-0.5B → Instruct (ours, 40 prompts, 5,223 tokens)unshifted: 86.7%marginal: 10.5%shifted: 2.8%86.7% · 10.5% · 2.8%Qwen2.5-0.5B → Instruct (ours, "Question:/Answer:" format)unshifted: 86.9%marginal: 11.1%shifted: 2.1%86.9% · 11.1% · 2.1%Llama-2-7b → Llama-2-7b-chat (URIAL)unshifted: 77.7%marginal: 14.5%shifted: 7.8%77.7% · 14.5% · 7.8%Llama-2-7b → Vicuna-7b-v1.5 (URIAL)unshifted: 82.4%marginal: 12.8%shifted: 4.8%82.4% · 12.8% · 4.8%Mistral-7b → Mistral-7b-instruct (URIAL)unshifted: 82.2%marginal: 12.5%shifted: 5.2%82.2% · 12.5% · 5.2%unshifted (base top-1)marginal (base rank 2-3)shifted (base rank > 3)
Unshifted, marginal and shifted tokens. Top two rows: our measurement on Qwen2.5-0.5B, first with the identical chat-template context, then with the base model reading a plain "Question: ... Answer:" context instead. Bottom three rows: the three 7B pairs reported by Lin et al. (2023), Figure 3. ch1_token_shift.py.
plain text
=== 40 prompts, 5223 answer tokens (format A: identical chat-template context) ===
unshifted (instruct token is base top-1): 86.7%
marginal  (base rank 2 or 3):              10.5%
shifted   (base rank > 3):                 2.8%
within base top-3:                         97.2%
format B (base sees "Question: ... Answer:"), 5200 tokens: unshifted 86.9%, within top-3 97.9%

The result reproduces, and then some. On 86.7% of the tokens of the instruct model's answers, the base model's own first choice was the very same token. On 97.2%, it was among the base model's top three. Only 2.8% of tokens are shifted. Lin et al. found 77.7% and 92.2% for Llama-2-7b-chat, a model 14 times larger from a different family trained with full RLHF, and about 82% unshifted for their two SFT-only pairs. Our pair agrees even more, which fits the point of Section 1.7: the Qwen2.5 base model has already seen a lot of assistant-style text.

We also checked that this is not an artefact of feeding the base model the chat template. In "format B", the base model reads Question: <q>\nAnswer: followed by the same answer text, with no special tokens at all. The numbers barely move (86.9% unshifted, 97.9% in the top three).

Which tokens shift

The script lists every shifted position with the tokens before it and the base model's preferred token. ch1_shift_analysis.py then sorts the 145 shifted tokens into groups:

Kind of tokenShifted tokensShare of shiftedShare of all answer tokensShift rate inside the group
end-of-turn <|im_end|>2416.6%0.5%100%
first token of the answer117.6%0.8%27.5%
punctuation, spacing, line breaks85.5%18.5%0.8%
other words10270.3%80.2%2.4%

The last column is the one to read. The end-of-turn token is shifted every time, the first token of an answer more than a quarter of the time, and an ordinary word only about one time in forty.

Four patterns stand out.

1. The end of the turn. The single most shifted token is <|im_end|>. It ends 24 of the 40 answers (the other 16 hit the 200-token limit), and it is shifted in all 24. The instruct model gives it an average probability of 0.81; the base model, on the same context, essentially zero (at most 0.00003). The base model's favourite at those positions was most often <|endoftext|>: it, too, thinks the text is over, but it uses the end-of-document marker from pretraining. This is the turn-taking behaviour of Section 1.8, measured.

2. The first token of the answer. How to open an answer is a style decision, and the two models often disagree. Of the 40 first tokens, 11 are shifted and 9 more are marginal. The base model's favourite openers are "The", "Sure", "To", "You" and "Hi"; the instruct model prefers "Certainly" (for the sleep tips and the Python function), "To" for maths ("To find..."), a quotation mark for a slogan or a bakery name, and "Dear" rather than "Hi" for a thank-you note. Notice that the base model's top choice was "Sure" on six prompts: it already knows how an assistant opens a reply, again a sign of the assistant-style data in its pretraining.

3. Formatting, less than you might expect. A few shifts are paragraph breaks (".\n\n" where the base model wanted "."). But punctuation and spacing tokens as a group shift less than 1% of the time, less than words do. Both models format answers in much the same way.

4. Word choice. Most shifted tokens (70%) are ordinary words, and many are synonyms or near-synonyms: " marks" where the base model wanted " occurs" (a sunset "marks" the end of the day), " display" for " sight", " engaging" for " clear", " Certainly" for " Sure". Creative writing and chat have the most of these, because there is no single right next word in a poem or a greeting.

And on the other side, look at what does not shift. The facts in the answers are unshifted: " Canberra" for the capital of Australia, " Leonardo" " da" " Vinci" for the Mona Lisa, " Aust" "en" for Pride and Prejudice. Qwen's tokenizer writes every digit as its own token, so the 206 bones of the human body are three tokens, "2" "0" "6", all unshifted. Over all answers, 245 tokens are digits (dates, quantities, results of sums): 98.8% of them are unshifted and none is shifted. The base model would have written the same facts.

Instruct answer to "Write a two-sentence story about a lost key.", coloured by how the base model ranked each tokenOne: base rank 37, KL 2.08, base top-1 SureOne day: base rank 1, KL 1.32, base top-1 day·day,: base rank 1, KL 0.03, base top-1 ,, a: base rank 1, KL 0.83, base top-1 a·a young: base rank 2, KL 0.86, base top-1 lost·young boy: base rank 3, KL 0.38, base top-1 man·boy stumbled: base rank 5, KL 0.57, base top-1 named·stumbled upon: base rank 1, KL 0.06, base top-1 upon·upon a: base rank 1, KL 0.07, base top-1 a·a dusty: base rank 4, KL 0.83, base top-1 lost·dusty old: base rank 1, KL 0.39, base top-1 old·old key: base rank 1, KL 0.12, base top-1 key·key in: base rank 1, KL 0.20, base top-1 in·in the: base rank 2, KL 0.00, base top-1 a·the attic: base rank 9, KL 0.81, base top-1 woods·attic of: base rank 1, KL 0.11, base top-1 of·of his: base rank 1, KL 0.03, base top-1 his·his grandmother: base rank 1, KL 0.37, base top-1 grandmother·grandmother's: base rank 1, KL 0.14, base top-1 's's house: base rank 1, KL 0.09, base top-1 house·house.: base rank 1, KL 0.15, base top-1 .. He: base rank 1, KL 0.76, base top-1 He·He had: base rank 2, KL 0.14, base top-1 was·had never: base rank 2, KL 0.13, base top-1 no·never seen: base rank 1, KL 0.15, base top-1 seen·seen such: base rank 2, KL 0.03, base top-1 it·such a: base rank 1, KL 0.01, base top-1 a·a key: base rank 1, KL 0.10, base top-1 key·key before: base rank 1, KL 0.03, base top-1 before·before,: base rank 1, KL 0.05, base top-1 ,, and: base rank 1, KL 0.08, base top-1 and·and it: base rank 2, KL 0.20, base top-1 he·it was: base rank 1, KL 0.12, base top-1 was·was a: base rank 1, KL 0.14, base top-1 a·a mystery: base rank 1, KL 0.19, base top-1 mystery·mystery to: base rank 1, KL 0.15, base top-1 to·to him: base rank 1, KL 0.01, base top-1 him·him.: base rank 1, KL 0.26, base top-1 .. The: base rank 2, KL 1.23, base top-1 He·The key: base rank 1, KL 0.03, base top-1 key·key was: base rank 1, KL 0.11, base top-1 was·was a: base rank 2, KL 0.17, base top-1 made·a key: base rank 3, KL 0.17, base top-1 relic·key that: base rank 1, KL 0.04, base top-1 that·that had: base rank 1, KL 0.14, base top-1 had·had been: base rank 1, KL 0.05, base top-1 been·been lost: base rank 1, KL 0.06, base top-1 lost·lost for: base rank 1, KL 0.20, base top-1 for·for years: base rank 1, KL 0.09, base top-1 years·years,: base rank 1, KL 0.11, base top-1 ,, and: base rank 1, KL 0.05, base top-1 and·and he: base rank 1, KL 0.13, base top-1 he·he was: base rank 1, KL 0.14, base top-1 was·was determined: base rank 1, KL 0.37, base top-1 determined·determined to: base rank 1, KL 0.00, base top-1 to·to find: base rank 1, KL 0.03, base top-1 find·find it: base rank 1, KL 0.01, base top-1 it·it.: base rank 1, KL 0.52, base top-1 .. With: base rank 4, KL 1.42, base top-1 He·With his: base rank 1, KL 0.24, base top-1 his·his grandmother: base rank 1, KL 0.31, base top-1 grandmother·grandmother's: base rank 1, KL 0.02, base top-1 's's help: base rank 1, KL 0.12, base top-1 help·help,: base rank 1, KL 0.06, base top-1 ,,unshiftedmarginalshiftedfirst 64 tokens; hover a token for its base rank and KL
The beginning of one instruct answer, every token coloured by its class. Most tokens are ones the base model would also have chosen; the few shifted ones are the opener and some word choices. Hover a token to see its base rank. ch1_token_shift.py.

We should also be honest about what the list shows that the paper's summary does not. A few shifted tokens change what is said, not only how. Asked "How do I pick a lock?", the instruct model begins "Picking a lock can be a fun ..." where the base model wanted "a bit ...". Asked "Hello! How are you today?", it says "I'm just a large language model" where the base model wanted "a system". In "How do I stay focused while studying?" the list of tips itself differs ("Set" a goal where the base model wanted "Take" a break). These are small, but they are not style: post-training changed the model's view of what it is and what it should say. At 0.5B parameters, neither model is reliable on facts, and our 40 prompts are too few to find factual changes reliably.

KL divergence: how different are two distributions?

The rank tells us whether the models agree on the top choice. It does not tell us how much the whole distribution moved. For that we use the Kullback-Leibler divergence.

KL(P ∥ Q)=∑v=1VP(v)log⁡P(v)Q(v)\mathrm{KL}(P \,\|\, Q) = \sum_{v=1}^{V} P(v) \log \frac{P(v)}{Q(v)}

where:

  • PP and QQ are two probability distributions over the same vocabulary, here P=PinstructP = P_{\text{instruct}} and Q=PbaseQ = P_{\text{base}} at one position;
  • the sum runs over all VV tokens vv;
  • log⁡P(v)Q(v)\log \frac{P(v)}{Q(v)} is positive where PP gives a token more probability than QQ, and negative where it gives less;
  • each term is weighted by P(v)P(v), so tokens the instruct model considers likely count most.

This is the same KL that appeared in the RLHF objective of Section 1.6, where it keeps the trained model close to the starting one. Here we use it as a ruler.

Worked example. ch1_kl_example.py computes the KL at two real positions. The first is the very first token of the answer to "Give me three tips for sleeping better." (a style decision). The second is the token after "The capital of Australia is" in an answer about Australia (a fact).

style: first token of the tips answerP_instructP_baseCertainlyP_instruct 0.56660.567P_base 0.08840.088term p·log(p/q) = +1.052SureP_instruct 0.17610.176P_base 0.35330.353term p·log(p/q) = -0.1231P_instruct 0.11150.112P_base 0.07360.074term p·log(p/q) = +0.046KL = 0.980 natsfact: after "The capital of Australia is"P_instructP_base·CanberraP_instruct 0.96290.963P_base 0.89290.893term p·log(p/q) = +0.073·MelbourneP_instruct 0.01230.012P_base 0.02850.029term p·log(p/q) = -0.010·SydneyP_instruct 0.01230.012P_base 0.03850.038term p·log(p/q) = -0.014KL = 0.046 nats
KL divergence at two real positions, the three most likely instruct tokens at each. Left: how to open an answer; the instruct model moved most of its probability to "Certainly". Right: a fact; both models put about 90% or more on " Canberra". ch1_kl_example.py.

For the style position, the four biggest terms of the sum are:

TokenPinstructP_{\text{instruct}}PbaseP_{\text{base}}Plog⁡(P/Q)P \log(P/Q)
Certainly0.56660.0884+1.052
Sure0.17610.3533−0.123
10.11150.0736+0.046
Here0.02320.0583−0.021
all other 151,932 tokens+0.025
total0.980 nats

Check the first term by hand: 0.5666×log⁡(0.5666/0.0884)=0.5666×log⁡6.41=0.5666×1.858=1.0520.5666 \times \log(0.5666 / 0.0884) = 0.5666 \times \log 6.41 = 0.5666 \times 1.858 = 1.052. The instruct model made "Certainly" 6.4 times more likely than the base model did, and since it gives that token 57% of its probability, this one term dominates. "Sure" contributes a negative term because the instruct model gives it less probability than the base model (the total is still positive; KL always is).

For the fact position, the KL is only 0.046 nats. The base model already put 0.893 on " Canberra"; the instruct model puts 0.963. The two models agree, with the instruct model a little more confident. Twenty times less divergence than at the style position.

Across all 5,223 tokens

plain text
KL(P_instruct || P_base) per token: mean 0.216 nats, median 0.041, 90th percentile 0.323, max 28.56
  mean KL at unshifted positions: 0.085  (4530 tokens)
  mean KL at marginal  positions: 0.371  (548 tokens)
  mean KL at shifted   positions: 3.743  (145 tokens)
  shifted positions are 2.8% of tokens but carry 48.1% of the total KL

The median token has a KL of 0.04 nats, about the same as the Canberra example: the two models have nearly the same distribution. The mean is five times larger than the median because a few positions are wildly different: the largest values (up to 28.6 nats) are at <|im_end|>, where the instruct model is sure and the base model gives almost nothing. The 2.8% of shifted positions carry almost half (48%) of all the divergence. Post-training changed a few positions a lot and most positions hardly at all.

The change is also concentrated at the start of answers:

Shift is largest at the start of the answer and fadesshifted tokens (%)mean KL per token (nats)tokens 1-5 (n=200)9.0%9.0%0.6350.635tokens 6-20 (n=555)5.0%5.0%0.4240.424tokens 21-50 (n=983)2.6%2.6%0.2160.216tokens 51-100 (n=1,453)2.8%2.8%0.1480.148tokens 101-200 (n=2,032)1.6%1.6%0.1670.167Includes the end-of-turn token <|im_end|>, which is always shifted and sits late in long answers.
Shifted tokens and mean KL by position in the answer, over all 40 answers. The first five tokens are shifted three to six times as often as tokens after the 50th, and their mean KL is about four times higher. ch1_token_shift.py.

Once the answer is under way, the instruct model's own earlier tokens are in the context, and the base model, reading them, is pulled along into the same style. Lin et al. report the same decline (their Figure 4). It is also why the URIAL paper's main proposal works: if you give a base model a good start, with a few carefully written example answers in its prompt (URIAL: "Untuned LLMs with Restyled In-context ALignment", with as few as three examples and a system prompt), it can answer like a chat model with no training at all.

The kind of prompt matters too. ch1_shift_analysis.py splits the result by the six groups of prompts:

Kind of promptAnswer tokensUnshiftedShiftedMean KL (nats)
knowledge1,04186.2%3.4%0.271
how-to, advice1,40786.4%2.3%0.130
creative writing66379.5%4.8%0.318
maths72694.5%1.0%0.214
code80089.8%1.2%0.108
chat, identity58683.1%4.8%0.361

Maths and code shift least: there is usually one right next token in a calculation or a line of code, and the base model already knows it. Creative writing and chat shift most: there are many good next words, and post-training has taught the instruct model its own preferences among them. These are small groups (about 600 to 1,400 tokens each), so treat the differences as indications, not precise measurements.

What the experiment shows, and what it does not

What it shows: for this pair of models, on these prompts, post-training moved the chosen token out of the base model's top three at fewer than 1 position in 30, and the changes are the end of the turn (always), the opening of the answer (often), and word choice here and there, most of all in creative writing. The factual tokens are the same ones the base model would have produced. This is strong support for the superficial alignment hypothesis, measured on our own models, and it matches the 7B results of Lin et al.

What it does not show:

  • It measures agreement on the instruct model's path. At every position the base model is given the instruct model's previous tokens. It is never allowed to go its own way. The base model's first choice differs from the instruct model's at 13% of positions (all the marginal and shifted ones), so a base model decoding on its own leaves the instruct model's path within a few tokens, and then the contexts are no longer the same. Section 1.7 showed where it can end up. Small local differences add up to a large difference in behaviour.
  • A few tokens can matter a lot. Refusing a harmful request, saying "I don't know", or stopping at the right moment are single tokens. 3% of tokens is not 3% of the value.
  • Greedy answers to easy prompts. Our 40 prompts are short and ordinary. Hard reasoning, long conversations, tool use and safety-critical prompts are where RL stages make their biggest changes, and they are not tested here. Section 1.6 noted that RLVR changes reasoning substantially; this experiment cannot see that.
  • Qwen2.5 base is not a pure base model. Its pretraining included synthetic assistant-written data (Section 1.6), so part of "alignment" may have happened before post-training even began. Our 86.7% may be higher than it would be for a model pretrained only on human-written text.
  • One small pair, 40 prompts, one decoding method. The numbers have sampling noise; another 40 prompts would give somewhat different percentages, and the repetition-penalty warning above shows that a small change in decoding moves them by several points.

"Superficial" is not "unimportant". The superficial part is what makes the model usable. But the evidence says that what post-training mostly does is select a style the base model already had available, rather than teach it new things to say.

1.10 A tiny training loop: SFT by hand

Everything so far has been measurement. Now we change some weights. The script ch1_sft_tiny.py takes the base model, Qwen2.5-0.5B, and fine-tunes all 494 million weights on 8 short chat examples for 30 steps. It is the SFT stage of Section 1.6 at the smallest possible scale, written out by hand so that every line is visible. It runs in a few minutes on a laptop GPU.

The data, and the masking that makes it SFT

Eight prompt-and-answer pairs such as ("What is the capital of Japan?", "The capital of Japan is Tokyo.") and ("Give me one tip for staying hydrated.", "Carry a water bottle and sip from it throughout the day."), plus three held-out pairs the model never trains on, to see whether what it learns carries over. Each pair becomes one token sequence and one list of labels:

python
def encode(q, a):
    msgs = [{'role': 'system', 'content': 'You are a helpful assistant.'}, {'role': 'user', 'content': q}]
    p = tok.apply_chat_template(msgs, add_generation_prompt=True, tokenize=True)   # prompt, ends with "<|im_start|>assistant\n"
    r = tok(a)['input_ids'] + [END]                                               # answer tokens + <|im_end|>
    return p + r, [-100] * len(p) + r                                             # input ids, labels

Line by line: the prompt is built with the chat template, exactly as at inference time, and ends by opening the assistant's turn. The response is the answer's tokens followed by <|im_end|>, the token that ends the turn. The input is the two joined together. The labels are the same tokens, except that every prompt position is replaced by −100. PyTorch's cross-entropy function skips any position whose label is −100. So the model reads the whole prompt, but is only graded on the answer and the closing <|im_end|>. That one line is the difference between the SFT loss and the pretraining loss of Section 1.6.

One SFT example: which tokens count in the loss<|im_start|>−100system−100↵−100You−100·are−100·a−100·helpful−100·assistant−100.−100<|im_end|>−100↵−100<|im_start|>−100user−100↵−100What−100·is−100·the−100·capital−100·of−100·Japan−100?−100<|im_end|>−100↵−100<|im_start|>−100assistant−100↵−100Theid·capitalid·ofid·Japanid·isid·Tokyoid.id<|im_end|>idprompt: label −100, ignored by the loss (26 tokens)answer + <|im_end|>: trained on (8 tokens)
The first training example, token by token. Grey tokens (the system prompt, the user turn and the assistant header) have label −100 and do not count. Only the 7 answer tokens and <|im_end|> are trained on. ch1_sft_tiny.py.

Of the 34 tokens in this example, only 8 are trained on. The script prints the labels so you can check:

plain text
  '<|im_start|>'   label -100
  'system'         label -100
  ...
  'assistant'      label -100
  '\n'             label -100
  'The'            label = 785
  ' capital'       label = 6722
  ' of'            label = 315
  ' Japan'         label = 6323
  ' is'            label = 374
  ' Tokyo'         label = 26194
  '.'              label = 13
  '<|im_end|>'     label = 151645
34 tokens, 8 of them are trained on

Why mask the prompt? Three reasons. First, we want to teach the model to answer, not to write questions; training on the prompt would spend effort predicting the user's words. Second, the system prompt and template are identical in every example, so without masking a large share of the loss would be spent on the same easy tokens. Third, it matches what the model will be asked to do at inference time: continue after <|im_start|>assistant\n.

The eight examples have different lengths, so they are padded to the same length with <|endoftext|>, and the padding gets label −100 too, plus an attention mask of 0 so that no real token pays attention to it.

The loop

1 · forwardrun the batch throughthe model → logits2 · losscross-entropy on theunmasked labels3 · backwarda gradient for eachof the 494M weights4 · updateAdamW moves eachweight a littlezero the gradients, take the next batch, repeatout = model(input_ids, labels=labels) → out.loss.backward() → opt.step(); opt.zero_grad()
One training step: forward pass, loss, backward pass, optimizer update. Then the gradients are cleared and the loop repeats.
python
opt = torch.optim.AdamW(model.parameters(), lr=1e-5, weight_decay=0.0)
model.train()
for step in range(1, 31):
    out = model(input_ids=ids, attention_mask=att, labels=lab)   # 1. forward pass, 2. loss
    out.loss.backward()                                          # 3. backward pass: gradients
    opt.step()                                                   # 4. update the weights
    opt.zero_grad()                                              #    clear the gradients

(simplified from ch1_sft_tiny.py, which also evaluates the held-out examples after every step). Every line, in order:

  • torch.optim.AdamW(model.parameters(), lr=1e-5) creates the optimizer, the rule that changes the weights. It is told which numbers it may change (all of them) and the learning rate, the size of each step. 1e-5 (0.00001) is a typical learning rate for fine-tuning a whole model; Qwen2.5's own SFT used 7e-6 decaying to 7e-7 (Qwen Team, 2024, Section 4.1).
  • model.train() switches the model into training mode (for models with dropout this turns dropout on; Qwen2.5 has none, but it is a good habit).
  • model(input_ids=..., attention_mask=..., labels=...) is the forward pass: it computes the logits at every position of all 8 sequences, then the cross-entropy loss of Section 1.5, averaged over every position whose label is not −100. out.loss is one number.
  • out.loss.backward() is the backward pass. PyTorch works backwards through every operation of the forward pass and computes, for each of the 494 million weights, the gradient: how much the loss would go up if that weight went up a little. This is backpropagation.
  • opt.step() changes every weight a little in the direction that makes the loss go down. AdamW scales each weight's step using running averages of its past gradients, which makes training much more stable than plain gradient descent.
  • opt.zero_grad() resets the gradients to zero. PyTorch adds new gradients to old ones by default, so without this line each step would use the sum of all previous gradients.

The output

plain text
step  0: train loss    -     held-out loss 2.252   P(<|im_end|>) on held-out 0.0000
step  1: train loss 2.442   held-out loss 1.691   P(<|im_end|>) on held-out 0.0001
step  2: train loss 1.791   held-out loss 1.409   P(<|im_end|>) on held-out 0.0001
step  3: train loss 1.248   held-out loss 1.387   P(<|im_end|>) on held-out 0.0001
step  4: train loss 1.035   held-out loss 1.342   P(<|im_end|>) on held-out 0.0001
step  5: train loss 0.894   held-out loss 1.342   P(<|im_end|>) on held-out 0.0002
step 10: train loss 0.784   held-out loss 1.566   P(<|im_end|>) on held-out 0.0003
step 20: train loss 0.683   held-out loss 1.642   P(<|im_end|>) on held-out 0.0009
step 30: train loss 0.575   held-out loss 1.593   P(<|im_end|>) on held-out 0.0027

held-out prompt: 'What is the capital of Italy?'
  before (stopped): 'The capital of Italy is Rome. Roma is the largest city in Italy and is also the capital of the country.<|endoftext|>'
  after  (stopped): 'The capital of Italy is Rome.<|im_end|>'
held-out prompt: 'How many legs does a spider have?'
  before (stopped): "A spider has 8 legs.\n戥user\nThat's correct! A spider has 8 legs.<|endoftext|>"
  after  (stopped): 'A spider has 8 legs.<|im_end|>'
held-out prompt: 'Give me one tip for better sleep.'
  before (did not stop): 'Sure, here are some tips for better sleep:\n\n1. Establish a regular sleep schedule: ...'
  after  (stopped): 'Carry a water bottle and sip from it throughout the day.<|im_end|>'

The "held-out loss" is the same SFT loss computed on the three examples the model never trains on, and "P(<|im_end|>)" is the probability the model gives to the end-of-turn token right after each held-out answer, averaged over the three.

loss, lr 1e-501230102030training steptrain 0.58held-out 1.59loss, lr 1e-401230102030training steptrain 0.00held-out 2.30P(<|im_end|>), held-out00.510102030training steplr 1e-5: 0.0027lr 1e-4: 0.9985
Two tiny SFT runs. Left and middle: training loss on the 8 examples (falls) and held-out loss on 3 new examples (falls, then rises: overfitting), at learning rates 1e-5 and 1e-4. Right: the probability of <|im_end|> after the held-out answers; only the larger learning rate learns it. ch1_sft_tiny.py.

Read the log from top to bottom.

The training loss goes down, as it must. From 2.442 at step 1 to 0.575 at step 30. Gradient descent on 8 fixed examples will always make their loss smaller.

The held-out loss goes down, then up. It falls from 2.252 to 1.342 in four steps: the model is learning something general about the format (answer briefly, in the assistant's turn). After that it climbs back to 1.593. The model has started to learn these 8 particular answers rather than the general habit. This is overfitting, and with 8 examples it starts almost at once.

The answers change shape. Before training, in the chat template, the base model answered "Rome" and then kept talking, or wrote a fake user turn ("戥user That's correct!"), and ended with <|endoftext|>, never <|im_end|>. After 30 steps, all three answers are one short sentence ending in <|im_end|>. The style of the 8 examples (one sentence, then stop) has been copied.

But one answer is the wrong answer. Asked for one tip for better sleep, the trained model replies "Carry a water bottle and sip from it throughout the day.", word for word the training answer for staying hydrated. That is memorisation in plain sight: the prompt "Give me one tip for ..." has been tied to one specific answer.

And it stops only by a hair. The log says the probability of <|im_end|> after the held-out answers is only 0.0027. How can greedy decoding pick a token with probability 0.27%? The script's check prints the top five tokens at that position:

plain text
  top-5 after the answer: '<|im_end|>' 0.0027, '<|im_start|>' 0.0010, 'oubted' 0.0006, 'powiedzieć' 0.0005, '看查看' 0.0005   (entropy 8.72 nats)

The distribution is almost flat: an entropy of 8.72 nats means the model is as unsure as if it were choosing among about e8.72≈6,000e^{8.72} \approx 6{,}000 tokens. <|im_end|> is the most likely, but only just. With sampling instead of greedy decoding, this model would almost never stop.

The reason is in the base model's weights. The script also measures the rows of the output table for <|im_start|>, <|im_end|> and the 290 unused rows after them:

plain text
base model output rows from <|im_start|> (151644) to 151935: 292 rows, cosine similarity of <|im_end|> to the others: min 1.0000, mean 1.0000

A cosine similarity of 1 means the vectors point in exactly the same direction. In the base model, the <|im_end|> row is indistinguishable from 291 rows that are never used. The base model never trained it, so it is not a "stop" token yet; it is a placeholder. SFT has to build that meaning from scratch, and 30 small steps are not enough.

A larger learning rate. Running the same script with python ch1_sft_tiny.py 1e-4 (ten times larger steps) shows the other side of the trade-off:

plain text
step 10: train loss 0.344   held-out loss 1.677   P(<|im_end|>) on held-out 0.0646
step 15: train loss 0.034   held-out loss 1.805   P(<|im_end|>) on held-out 0.8395
step 30: train loss 0.000   held-out loss 2.298   P(<|im_end|>) on held-out 0.9985

held-out prompt: 'How many legs does a spider have?'
  after  (stopped): 'Carry a spider and sip from it throughout the day.<|im_end|>'

Now <|im_end|> is learned properly: probability 0.9985, higher than the instruct model's 0.81 average from Section 1.9. But the training loss is 0.000 (every training answer is memorised exactly), the held-out loss has risen to 2.298, above where it started, and the spider question now gets "Carry a spider and sip from it throughout the day." The model learned when to stop and forgot how to answer.

This is the whole problem of SFT in miniature. The format, including the end-of-turn token, has to be learned, and that takes enough training. But with too few and too similar examples, the same training memorises instead of generalising. Real SFT solves it with data: hundreds of thousands of varied examples (Qwen2.5 used over a million), so that the only thing common to all of them is the format and the helpful style. LIMA's 1,000 examples worked because they were carefully varied, and because the base model was 130 times larger than ours. A later part of the book covers SFT data, learning-rate schedules and how to tell when to stop.

1.11 How do we know it got better? A first look

Every stage of the pipeline claims to make the model better. How would you check? Chapter 3 is about evaluation in depth; here is a first taste, with a tiny test anyone can run.

The script ch1_eval_preview.py asks 20 short questions that each have one clear answer: four capitals, two authors and scientists, six arithmetic questions, and some general knowledge ("What is the largest ocean on Earth?", "What is the longest river in Africa?"). For each model setting it checks two things: does the expected answer (for example "Nairobi", "72", "Pacific") appear anywhere in the output, and did the model end its turn by itself within 60 new tokens? Greedy decoding again.

20 short questions, greedy, at most 60 new tokensexpected answer appearsended its turn by itselfbase, raw question12/2012/209/209/20base, chat template14/2014/205/205/20instruct, chat template18/2018/2016/2016/20
A first evaluation on 20 short questions. Left: how often the expected answer appears in the first 60 new tokens. Right: how often the model ended its turn within those 60 tokens. ch1_eval_preview.py.
plain text
base, raw question       answer found: 12/20   ended its turn by itself:  9/20
base, chat template      answer found: 14/20   ended its turn by itself:  5/20
instruct, chat template  answer found: 18/20   ended its turn by itself: 16/20

At first sight the instruct model is much better: 18 of 20 against 12, and it ends its turn on 16 of 20 questions against 9. But read the misses before you believe a score. The script prints every one:

  • All six arithmetic misses of the raw base model are cut-offs, not wrong answers. For "What is 9 times 8?" it begins "To find the product of 9 and 8, we can use the standard multiplication method. Here are t..." and runs out of its 60-token budget before reaching 72. With a larger budget it would probably have scored more. The test, not only the model, decided that result.
  • Two raw base misses are "wrong genre". "Who wrote Romeo and Juliet?" is continued with more questions ("What is the name of the play? What is the name of the author?..."), and "How many continents are there?" becomes a multiple-choice exam ("A. 1 B. 2 C. 3 D. 4 Answer: C"). The knowledge may be there; the format is not.
  • The base model in the chat template fails differently: for four questions it falls straight into a loop of junk tokens (" Comey", " Cherokee", "ounces" repeated), and once it states a wrong fact ("plants take in oxygen from the air").
  • The instruct model's two misses are real errors, stated confidently. "There are currently 5 continents: Asia, Africa, North America, South America, and Europe." and "There are 60 minutes in two hours." Post-training made the answers short, polite and well-formed; it did not make a 0.5B model right. The base model, reading the same question as plain text, at least started to work out the minutes step by step.
  • "Ended its turn" is also budget-bound. The four instruct answers that did not stop within 60 tokens were explanations that kept going (about relativity, the Pacific, photosynthesis and the Nile). Whether that counts as a failure depends on whether you wanted a one-word answer.

Four lessons, which Chapter 3 develops properly:

  1. A score is a measurement procedure, not a property of the model. The prompt format, the token budget and the matching rule changed our numbers as much as the model did.
  2. Read the outputs. A substring check counts "Answer: C" as a miss and would count "not 72" as a hit. Every automatic metric needs a look at real examples.
  3. Measure what you care about. If you want an assistant, measure answers in the assistant format and also how it behaves (stopping, length, refusals), not only whether a fact appears.
  4. Twenty questions is a demonstration, not an evaluation. With 20 items, one question is 5 percentage points. Real benchmarks use thousands of items, careful answer checking, and guard against test questions having leaked into the training data.

The published reports do this at scale. The InstructGPT result quoted in Section 1.6 (labelers preferred the 175B InstructGPT 85% of the time) comes from human comparisons of outputs on held-out prompts; the Qwen2.5 and Llama 3 reports each list dozens of benchmarks for knowledge, maths, code and instruction following before and after post-training. Chapter 3 explains how such numbers are made and how far to trust them.

Takeaways

References

Papers

  • Brown, T. et al. (2020). Language Models are Few-Shot Learners (GPT-3). arXiv:2005.14165
  • Chung, H. W. et al. (2022). Scaling Instruction-Finetuned Language Models (Flan-T5, Flan-PaLM). arXiv:2210.11416
  • DeepSeek-AI (Guo, D. et al.) (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948
  • Grattafiori, A. et al. (2024). The Llama 3 Herd of Models. arXiv:2407.21783
  • Lambert, N. et al. (2024). Tulu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv:2411.15124
  • Lin, B. Y. et al. (2023). The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning (URIAL). arXiv:2312.01552
  • Ouyang, L. et al. (2022). Training language models to follow instructions with human feedback (InstructGPT). arXiv:2203.02155
  • Qwen Team (2024). Qwen2.5 Technical Report. arXiv:2412.15115
  • Rafailov, R. et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290
  • Shao, Z. et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (introduces GRPO). arXiv:2402.03300
  • Team OLMo (Walsh, P. et al.) (2024). 2 OLMo 2 Furious. arXiv:2501.00656
  • Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762
  • Wei, J. et al. (2021). Finetuned Language Models Are Zero-Shot Learners (FLAN). arXiv:2109.01652
  • Zhou, C. et al. (2023). LIMA: Less Is More for Alignment. arXiv:2305.11206

Other sources