Speculative Decoding: Several Tokens per Step, Not One Word Changed
How a small model can guess ahead and a big model can check all the guesses at once, why the output stays exactly the same, and what it really buys on real hardware: speculative decoding built from scratch and measured, from 1.16x on prose to 1.89x on code and 3.49x when the answer copies the prompt.
Part 1 left us with a strange fact. On our test machine, pushing 16 tokens through a model in one pass took 10.8 ms. Pushing one token through took 10.0 ms. Sixteen times the work, for almost the same time.
That is decode's big inefficiency in one line. Every step, the GPU hauls all of the model's weights out of memory (the expensive part) to do a tiny amount of math with them (the cheap part). The compute sits mostly idle.
This part is about a clever way to put that idle compute to work. It lets a model produce several tokens per step instead of one, and it does it without changing a single word of the output. It is called speculative decoding, and by the end you will have built it from scratch and seen exactly when it pays off and when it does not.
First: checking is cheap, writing is not
Why would checking guesses be cheaper than writing? Go back to the France example from Part 1. Writing "The capital is Paris." took four decode steps, one token each, because each word depended on the one before.
But suppose someone handed you the finished sentence and asked, "is this what you would have written?" Now every word is known. The model can read the whole sentence at once, exactly like a prefill, and at every position ask: "given everything before this word, is this the word I would have picked?" One pass answers the question for all of them.
That asymmetry is the whole trick. Producing tokens is sequential. Checking tokens is parallel.
Here it is measured on the model we will use as our "big" model in this part, Qwen2.5-3B-Instruct:
| Tokens checked in one pass | Time |
|---|---|
| 1 | 30.9 ms |
| 5 | 33.7 ms |
| 9 | 34.4 ms |
| 17 | 38.2 ms |
Checking 17 tokens takes only about 1.2 times as long as producing one. The weights get read once either way.
Draft, then verify
So we need someone to make guesses. The classic choice is a draft model: a much smaller model from the same family that uses exactly the same vocabulary of tokens (it has to: the two models' tokens are compared one by one). Here that is Qwen2.5-0.5B-Instruct, six times smaller than the 3B target. It is not as smart, but on easy stretches of text it usually guesses what the big model would say.
Each round of speculative decoding has three steps:
- Draft. The small model writes the next K tokens the normal way, one at a time. It is small, so each of those steps is quick.
- Verify. The big model takes all K guesses and runs one pass over them, computing at every position which token it would have chosen.
- Accept. Walk along the guesses. Keep each one that matches what the big model would have picked. At the first mismatch, stop, throw away that guess and everything after it, and use the big model's choice instead.
Notice step 3 always produces at least one token. Even if the very first guess is wrong, the big model has already computed what the right token is, so the round still moves forward by one. Speculative decoding can never produce fewer tokens per big-model pass than plain decoding. It can only waste the draft model's time.
A real round, from our run
Here are the first few rounds of an actual run, with K = 4 guesses per round, on the prompt "Explain how a refrigerator keeps food cold":
| Round | Draft guessed | Kept | Big model adds | Tokens this round |
|---|---|---|---|---|
| 1 | refrigerator works by utilizing | 1 of 4 | keeps | 2 |
| 2 | food cold by utilizing | 2 of 4 | through | 3 |
| 3 | a process called refriger | 4 of 4 | ation | 5 |
| 4 | . Here 's a | 1 of 4 | It | 2 |
| 5 | works by using a | 4 of 4 | cycle | 5 |
| 6 | of heat and cold | 1 of 4 | heating | 2 |
| 7 | and cooling to transfer | 2 of 4 | . | 3 |
| 8 | When food is placed | 0 of 4 | The | 1 |
Round 3 is the dream case: the draft guessed a process called refriger, the big model agreed with all four, and added ation itself. Five tokens for one pass. Round 8 is the worst case: the first guess was wrong, so the round produced just the big model's own token, exactly what plain decoding would have produced, and only the draft's time was lost.
Every round that keeps even one guess is a round where the big model did more than one token's worth of progress for the price of one pass.
Why the output does not change at all
This is the part that makes speculative decoding special. It is not an approximation. Done right, the text that comes out is exactly what the big model would have produced on its own.
With greedy decoding (always pick the most likely token) it is easy to see why. We only keep a guess if it is the same token the big model would have picked at that position. And when a guess is wrong, we take the big model's own pick. Every token that reaches the output is a token the big model chose. The draft model only decides how many of them we get per pass.
With sampling (picking randomly according to the model's probabilities, which is how most chatbots run) it needs one more idea, from the two papers that introduced the method, Leviathan et al. and Chen et al. (both 2023). Let be the big model's probability for token , and the draft model's. The draft proposes . Then:
- keep it with probability ;
- if it is rejected, draw a replacement from what the draft under-weighted: the distribution proportional to .
If the big model likes a token at least as much as the draft did, it is always kept. If the draft was over-eager about a token, it is kept only some of the time, and the replacement step adds back exactly the probability the draft missed. The two steps together rebuild exactly.
You do not have to take that on faith. Here is a simulation of one million tokens with a draft that is badly miscalibrated: it thinks "a" is the most likely word, when the big model prefers "the".
The output matches the big model's distribution to within 0.0006 on every word. And the share of guesses that were kept came out at 72.0%, exactly the theory's prediction of . A worse draft does not make the output wrong. It only makes the process slower.
The economics: how many tokens per pass?
How much faster this is depends on one number above all: the acceptance rate , the chance that a given guess is right. The trouble is that acceptance compounds. Guess 3 only counts if guesses 1 and 2 were right too.
If each guess is right with probability and you make guesses, the expected number of tokens per big-model pass is
(Leviathan et al., 2023). Here is what that looks like:
Two lessons fall out of this curve:
- Acceptance matters more than anything. At 90%, eight guesses give about 6 tokens per pass. At 50%, you never get past 2, no matter how many guesses you make.
- Guessing further has diminishing returns, and every extra guess still costs a draft step. So there is a best K, and it depends on how predictable the text is.
The full cost of a round is K draft steps plus one verify pass. Speculative decoding wins only if the tokens you gain outweigh the draft steps you spend.
Measured: what it really buys
I implemented speculative decoding from scratch (the code is at the end), with Qwen2.5-3B-Instruct as the big model and Qwen2.5-0.5B-Instruct as the draft, greedy decoding, 200 tokens per answer, on an Apple M5 Pro GPU. Plain decoding with the 3B model ran at about 31 tokens per second. One draft step cost 9.9 ms and one target pass 30.9 ms.
| Task | Best K | Speedup | Guesses kept | Tokens per big-model pass |
|---|---|---|---|---|
| rewrite a paragraph (copy with small edits) | 8 | 1.86x | 79% | 6.7 |
| write a Python function | 8 | 1.89x | 81% | 7.1 |
| explain how a fridge works | 2 | 1.16x | 58% | 2.1 |
| write an original story | 1 | 1.14x | 67% | 1.7 |
The pattern is exactly what the economics predicted:
- Predictable text wins big. Code and the rewrite task kept 80 to 98% of the draft's guesses. Every extra guess kept paying off, all the way to K = 8, where the big model produced 7.1 tokens per pass on the code task, for a 1.89x speedup.
- Open-ended text barely wins. The explanation and the story kept only 25 to 65% of guesses. The best result was one or two guesses per round, for about 1.15x. With more guesses it got slower than plain decoding: at K = 8 the story ran at 0.72x, because most guesses were thrown away and each one still cost a draft step.
- The drafter is expensive here. One step of the 0.5B draft took 9.9 ms, about a third of a 3B pass (30.9 ms), because on this machine a small model's step is dominated by fixed overhead rather than by its size. A drafter that cheap relative to the target is the main reason the ceiling is under 2x. On a data-centre GPU serving a 70B model, where the big model's step is far more expensive than the drafter's, the same machinery has much more room: that is where the published 2x to 3x results come from.
Every run was also checked against plain decoding, token by token. More on that in a moment.
A drafter that costs nothing: prompt lookup
A draft model is not the only way to guess. In a lot of real work the answer copies from the prompt: summarising a document, answering questions about it, editing code, rewriting text. There, the best guess for "what comes next" is often "whatever came next the last time this phrase appeared".
Prompt lookup decoding (Saxena, 2023) does exactly that. Take the last few tokens written, search the prompt and the answer so far for the same sequence, and propose whatever followed it. No second model, no extra memory, and the search takes microseconds.
On the rewrite task, where the answer is the prompt's paragraph with a few words changed, prompt lookup was the fastest thing in this whole article: 3.49x with 8 guesses per round, and 2.70x with 4. It beat the draft model (1.86x) because its guesses cost nothing.
On everything else it did little or nothing: 1.23x on code, and slightly slower than plain decoding on the explanation and the story (0.93x and 0.97x), where there is almost nothing to copy and each failed lookup still means an extra check. It is the right tool when the output copies the input, and the wrong one otherwise.
Better drafters
Most of the research since 2023 has been about getting better guesses more cheaply:
- Medusa (Cai et al., 2024) skips the separate draft model. It adds a few extra prediction "heads" to the big model itself, each guessing a different position ahead, and checks several candidate continuations at once with a tree-shaped attention pattern. The paper reports over 2.2x speedup with the original model frozen.
- EAGLE (Li et al., 2024) trains a very small network that works on the big model's own internal features rather than on raw tokens, which makes its guesses far more accurate. The paper reports 2.7x to 3.5x lower latency on LLaMA2-Chat 70B. EAGLE-3 (2025) reports up to 6.5x.
- Multi-token prediction trains the model to predict several future tokens from the start, which gives it a built-in drafter.
vLLM supports all of these. A draft model looks like this:
vllm serve Qwen/Qwen2.5-3B-Instruct \
--speculative-config '{"method": "draft_model", "model": "Qwen/Qwen2.5-0.5B-Instruct", "num_speculative_tokens": 4}'and prompt lookup like this:
vllm serve Qwen/Qwen2.5-3B-Instruct \
--speculative-config '{"method": "ngram", "num_speculative_tokens": 4, "prompt_lookup_min": 2, "prompt_lookup_max": 3}'When it helps, and when it hurts
Everything above comes back to Part 1. Speculative decoding spends idle compute to save sequential steps. So it helps exactly when there is idle compute, and stops helping when there is not.
| It helps when | It hurts or does nothing when |
|---|---|
| the text is predictable (code, structured output, copying from the input) | the text is open-ended and creative, so acceptance is low |
| the GPU is lightly loaded, so each step is memory-bound with compute to spare | the GPU is already busy with a large batch (Part 1): the "spare" compute is no longer spare |
| the big model is slow relative to the drafter | the drafter is expensive relative to the big model |
| you care about latency for one user | you care only about total throughput at high load |
That is why vLLM's documentation lists the biggest gains at low load ("latency focused"), and why serving systems adjust speculation, or switch it off, as load grows (vLLM's documentation lists a dynamic speculative decoding option for exactly this). It is a tool for making one user's answer arrive faster, not a free throughput boost.
Summary
- Decode leaves the GPU's compute mostly idle. Checking tokens is parallel and costs about the same as producing one: 38.2 ms to check 17 tokens against 30.9 ms for one, on our 3B model.
- Speculative decoding: a cheap drafter guesses K tokens, the big model checks them in one pass, keeps the correct prefix and adds its own next token. Never fewer than one token per pass.
- It is exact. With greedy decoding, every kept token is the big model's own choice. With sampling, the accept-with-probability rule plus resampling from reproduces the big model's distribution, confirmed here to within 0.0006.
- The payoff depends on the acceptance rate: expected tokens per pass .
- Measured: a from-scratch implementation reached 1.89x on code and up to 3.49x with prompt lookup on a rewrite task, but only about 1.15x on open-ended prose, and it got slower than plain decoding with too many guesses.
- It is a latency tool: best at low load, on predictable text, with a cheap drafter.
Speculative decoding from scratch (greedy)
"""Greedy speculative decoding with a draft model, from scratch.
target: the big model, draft: a small model with the same tokenizer (e.g. Qwen2.5-3B and 0.5B)."""
import torch
from transformers import DynamicCache
@torch.inference_mode()
def run(model, ids, cache):
return model(input_ids=torch.tensor([ids], device=model.device), past_key_values=cache).logits[0]
@torch.inference_mode()
def speculative(target, draft, prompt, k=4, max_new=200, eos=None):
tcache, dcache = DynamicCache(), DynamicCache()
out = [int(run(target, prompt, tcache)[-1].argmax())] # prefill: the first token is the target's
run(draft, prompt, dcache)
while len(out) < max_new and out[-1] != eos:
committed = prompt + out # out[-1] is not in the target's cache yet
# 1. the draft guesses k tokens, one at a time
guesses, x = [], committed[dcache.get_seq_length():]
for _ in range(k):
g = int(run(draft, x, dcache)[-1].argmax())
guesses.append(g)
x = [g]
# 2. the target checks every guess in ONE pass: preds[i] is its choice after position i
preds = run(target, [out[-1]] + guesses, tcache).argmax(-1).tolist()
# 3. keep guesses while they match, then take the target's own token
n = 0
while n < k and guesses[n] == preds[n]:
n += 1
keep = len(prompt) + len(out) + n # roll both caches back past rejected guesses
tcache.crop(keep)
if dcache.get_seq_length() > keep:
dcache.crop(keep)
out += guesses[:n] + [preds[n]]
return out[:max_new]Plain greedy decoding with the target alone produces exactly the same tokens (up to floating-point ties), one per pass. The only difference is how many target passes it takes.
References
- Y. Leviathan, M. Kalman, Y. Matias. Fast Inference from Transformers via Speculative Decoding. ICML 2023. (2x to 3x on T5-XXL "with identical outputs.")
- C. Chen et al. Accelerating Large Language Model Decoding with Speculative Sampling. 2023. (2x to 2.5x on Chinchilla 70B, preserving the target distribution "within hardware numerics.")
- T. Cai et al. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads. 2024.
- Y. Li et al. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty. 2024.
- Y. Li et al. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test. 2025.
- A. Saxena. Prompt Lookup Decoding. 2023.
- vLLM documentation: Speculative decoding.