How Models Are Trained
From next-word predictor to assistant: pretraining, SFT, RLHF, DPO and RL for reasoning
How a language model goes from predicting the next word to following instructions, holding a conversation and reasoning step by step. The history from 2017 to today, the research papers as highlighted screenshots, every equation explained symbol by symbol, and real code run on real models, in plain English for beginners.
A model that has only read the internet can finish your sentence, but it cannot answer your question. Something happens between the raw model and the assistant you talk to, and that something has a history, a set of equations and a lot of practical tricks. This book explains all of it, one step at a time.
It is written for beginners, in simple English. Every idea comes with the research paper that introduced it (shown as highlighted screenshots), a simple picture, the maths explained one symbol at a time with real numbers, and small pieces of real code that run on a laptop, with every line explained.
Part 1, the big picture, covers what training actually changes inside a model, how the methods began and evolved, and how anyone can tell whether a model got better.
Chapters
- Chapter 1 · From next word to assistant
What a language model computes (tokens, next-token probabilities, the chain rule, softmax, cross-entropy, perplexity), the full training pipeline from pretraining to RL with verifiable rewards, and hands-on experiments with Qwen2.5-0.5B base and instruct: the chat template, how few tokens post-training really changes, and a tiny SFT loop.
- Chapter 2 · How it started, and how it evolved
The history of how language models learned to follow instructions, told as one story where each step fixes a problem left by the step before: from Shannon's n-grams and the Transformer, through pretraining, learning from human preferences and instruction tuning, to InstructGPT, ChatGPT, DPO, GRPO and DeepSeek-R1. With the papers' own figures, the four key equations worked on real numbers, and small experiments you can run.