Chapter 3 · Reasoning patterns: how one agent thinks, acts and corrects itself
Reasoning patterns for one agent: chain of thought, ReAct, plan-and-execute, reflection with an external check, tree search and code as action.
Goal: by the end of this chapter you can take one task, try it with several loop shapes (think, act, plan, check, choose among candidates), say what each shape changes, what it costs and how it fails, and decide with a measurement which one to keep. You will read the key papers behind each pattern, see three small experiments run against a real model, and have one complete modern tool loop you can copy.
Chapter 1 drew the agent as a loop and Chapter 2 took the loop apart into blocks. This chapter keeps the blocks and changes one thing: the shape of the loop. How many times is the model called? Does it write its reasoning down? Does it look things up as it goes, or plan everything first? Does anything check the answer? Does it try several answers and pick one?
We will follow one example the whole way through.
Two questions come back again and again:
How many years after the Porto depot opened did the Leeds depot open? The answer is 3 (2020 minus 2017). It needs two look-ups and a subtraction.
Which depot has the most trucks, and how many more does it have than the Leeds depot? The answer is Rotterdam, with 79 more (120 minus 41). It needs four independent look-ups, a comparison and a subtraction.
Here is what each shape does with the first question, before any theory:
Shape
What the system does with the Porto and Leeds question
What happened in our runs (gpt-5-mini)
Direct answer
one call, answer only
"6 years", "9 years", "5": invented every time
Chain of thought
one call, write the steps first
two of three runs said "I don't have the opening years"; one said "2 years"
ReAct loop
think, look up, observe, repeat
right in one run of three; two runs went round in circles until the step budget ran out
Loop + guard + better tool
the same loop, with two engineering fixes
right in three runs of three
Plan first
write all look-ups, run them, answer once
not run on this question; Section 3.4 counts its cost
Check the answer
recompute from the evidence before replying
catches a wrong number before the user sees it (Section 3.5)
The rest of the chapter explains each row: the idea, the paper that introduced it, what it costs, how it fails and what to measure.
Six shapes of the loop with the same model and tools: think only, act only, interleave (ReAct), plan first, reflect, and search.
The shapes are choices on five separate dimensions#
It is tempting to read the six shapes as a ladder, each one "more advanced" than the last. That is the wrong picture. Each pattern changes one or two separate design choices, and a real agent picks an option on each.
Five independent design dimensions of a single agent. A pattern is a choice in one or two rows; the highlighted cells are the structured loop of Section 3.3.
Dimension
The question it answers
Patterns that change it
Reasoning inside one call
how does the model think before it answers?
chain of thought, self-consistency, reasoning models
Control across calls
who decides the next step?
ReAct, plan-and-execute, Agentless
Action representation
what does an action look like?
function calling, CodeAct
Verification and selection
what decides that an answer is good?
reflection, CRITIC, best-of-n, tree search
Context management
what does each call see?
ReWOO, compaction, prompt caching
So "we use ReAct" says something about control and nothing about verification. Every section below says which row it changes.
Every pattern is judged on three numbers measured on a labelled set: accuracy, cost per task (tokens and tool calls) and latency (time to answer). The number that combines the first two is:
Most papers below used models that are now several generations old. Read their tables for the shape of each result (which pattern helps on which task, at what cost) and re-measure on your own model.
3.2 Thinking inside one call: chain of thought, voting and reasoning models#
Row changed: reasoning inside one call.
Start with the public question "How many years passed between the completion of the Eiffel Tower and the opening of the Sydney Opera House?" (1973 minus 1889, so 84). Asked for the answer only, gpt-5-mini said 66, 72 and 75 in three runs. Asked to "think step by step" first, it wrote down 1889 and 1973, subtracted, and said 84 all three times. Nothing else changed: the same model, the same question, one call each.
Now the Northwind question about Porto and Leeds. Thinking step by step did not help. In two runs the model said, correctly, that it did not have the opening years; in the third it wrote "2 years". On the question about the chief executive it invented three different names in three runs. A scratchpad helps a model use what it knows. It cannot create a fact the model never saw.
Standard prompting against chain of thought, and self-consistency: several sampled chains and a plurality vote over their final answers.
In January 2022 Jason Wei and colleagues at Google Research showed that a large enough model, given a few examples whose answers contain the working, writes working for new problems too and gets many more of them right. No training was involved.
On GSM8K (grade-school maths word problems) PaLM-540B went from about 18% to about 57%. The gain appeared only in large models; small models often got worse. Four months later Takeshi Kojima and colleagues showed that one sentence, "Let's think step by step", gets most of the gain with no examples.
If one chain is good, are several better? Xuezhi Wang and colleagues (2022) sampled many chains at a non-zero temperature, read off each final answer and returned the most common one. They called it self-consistency. With 40 samples, PaLM-540B on GSM8K rose from 56.5% to 74.4%.
Two details matter in practice.
Plurality, not majority. Suppose each chain is right 40% of the time, and wrong chains split evenly between two different wrong answers, 30% each. No answer has a majority, yet over many samples the right answer is the most frequent and wins. Voting needs the right answer to be more common than any single wrong answer, not more common than all wrong answers together.
Correlated errors. The argument above assumes the wrong chains scatter. If they make the same mistake, voting makes it more certain, not less. Our chain-of-thought runs show this: on Everest against Kilimanjaro all three chains answered 2,953 metres, because the model remembers Everest as 8,848 m while our encyclopaedia says 8,849 m. A vote over those three chains would return 2,953 with full agreement. Agreement measures consistency, not correctness.
Self-consistency was not the first result to spend extra computation at answer time. In October 2021 Karl Cobbe and colleagues at OpenAI generated many candidate solutions per problem and picked the one a trained verifier ranked highest (Section 3.6). Self-consistency replaced the trained verifier with agreement between samples.
From late 2024 the pattern moved into training. OpenAI's o1 series, then DeepSeek-R1 and others, were trained with reinforcement learning to produce a long chain of thought before answering, rewarded when the final answer could be checked.
What reasoning models changed about chain of thought and what they did not. Visibility and controls differ across models and providers.
For an agent builder, much stays the same: the model still cannot look anything up, still needs a stop condition and a budget, and is still wrong some of the time. What changed is the economics (with some providers you pay for reasoning you cannot read in full) and the question ("should I ask it to think?" becomes "how much effort should I buy for this step?"). Our experiments use gpt-5-mini with its effort set to minimal, so the reasoning we study is the visible kind.
This result was measured on 2023 models. Section 3.8 shows the 2025 follow-up on reasoning models: the chains became more faithful but were still far from fully faithful in the settings studied. The practical rule is conditional: a chain is useful for debugging, but where correctness matters, do not treat a plausible chain as evidence. Check the answer against something outside the text.
What to measure. Accuracy on the multi-step slice of your labelled set with and without a chain, output tokens per call, and, for a reasoning model, accuracy gained per dollar at each effort setting.
Research notes: the chain-of-thought papers in detail (optional reading)
These are the remaining excerpts and table readings behind Section 3.2. Nothing here is needed for the rest of the chapter.
Row changed: control across calls (and, once tools appear, the action format).
Thinking cannot produce the Porto depot's opening year. Looking it up can. The simplest loop that looks things up is: think about what you need, call a tool, read the result, repeat. Here is what that loop did on the Porto and Leeds question when it worked:
plain text
Thought: I need the opening years of the Porto and Leeds depots of Northwind Freight.Action: wiki[Northwind Freight]Observation: Fictional logistics company founded in 2011 in Rotterdam. ... Related entries: Rotterdam depot, Gdansk depot, Porto depot, Leeds depot, ...Action: wiki[Porto depot] Observation: Northwind Freight depot opened 2017. Vehicles: 64 trucks. ...Action: wiki[Leeds depot] Observation: Northwind Freight depot opened 2020. Vehicles: 41 trucks. ...Action: calc[2020-2017] Observation: 3Action: finish[3]
And here is the same plain loop on a public question, when it did not:
After step 3 the answer, 84, is in the context. The model computes it again, then starts the look-ups over, and the step budget runs out. This failure has a name in the paper that introduced the loop.
The ReAct loop and its two surface forms: the 2022 text format parsed by a regular expression, and the same loop with structured tool calls.
Shunyu Yao (Princeton) and colleagues at Google posted ReAct in October 2022. Chain of thought could reason but could not look anything up, so it invented facts. Earlier acting agents could call tools but kept no reasoning between actions, so they lost track of the goal. ReAct put both in one loop. For question answering the tools were three operations over Wikipedia (search, lookup, finish) and the budget was seven steps.
Two readings matter. First, ReAct lost to chain of thought on HotpotQA (27.4 against 29.4); the best results combined the two with a fallback. Second, the paper's hand-labelled error study (Table 2, read in the research notes below) found the two failures our trace shows: "reasoning error", which the authors define to include failing to recover from repetitive steps (47% of the sampled ReAct failures), and non-informative searches that derail the run (23%). The same table is often quoted as "ReAct has 0% hallucination". That is too strong. The 0% is the hallucination category among sampled failed ReAct trajectories; among sampled successful ones, 6% were right with a hallucinated reasoning trace or facts. Grounding reduced hallucination in that study; it did not remove it.
The script ch3_react.py runs the twelve questions (six public, six about Northwind) through six configurations with gpt-5-mini at minimal reasoning effort, three times each, so every rate is over 36 runs. The four loop configurations isolate two engineering fixes:
Repeated-action guard (in the loop): if the model emits an action it has already run, the loop replaces the observation with "you have been round this loop; using only the observations above, finish now".
Improved tool (in the tool): an entry lists related entries by name (the company entry names its four depots as "Porto depot" and so on), several matches come back as entries instead of an error, and a miss explains how entries are named.
Grading is typed and exact. An answer is right only if it contains exactly the expected numbers, compared as numbers, and names every expected entity after explicit normalisation (accents removed, lower case, punctuation collapsed). An earlier draft of this script used substring matching, which marked "184" correct for 84. The script now starts with a self-test that shows the difference:
python
# code/agents/ch3_react.py (abridged): the typed grader and the guarddef grade(answer, expected): got = sorted(numbers_in(answer)) # "2,954 m" -> [2954.0]; "three" -> [3.0] want = sorted(float(x) for x in expected.get('numbers', [])) if len(got) != len(want) or any(abs(g - w) > 1e-9 for g, w in zip(got, want)): return False # "184", "13", "It is not 84; the answer is 74." all fail text = f' {normalise(answer)} ' return all(f' {normalise(e)} ' in text for e in expected.get('entities', []))if (tool, arg) in seen: # inside the loop repeats += 1 if guard: obs = 'You have been round this loop once. Using ONLY the observations above, ... finish[answer] now.'
Six system configurations, 36 runs each: the guard and the improved tool each fix different questions; only both together answer all twelve.
Configuration
Accuracy (36 runs)
Public
Private
Calls
Cost per correct
Runs that hit the budget
Direct answer
3%
6%
0%
36
$0.00144
0
Chain of thought
22%
44%
0%
36
$0.00086
0
ReAct, original tool, no guard
56%
44%
67%
189
$0.00172
16
+ guard only
75%
100%
50%
134
$0.00085
0
+ improved tool only
75%
56%
94%
170
$0.00111
9
+ guard + improved tool
100%
100%
100%
150
$0.00072
0
Read it in three steps.
Direct against chain of thought: a scratchpad helps where the facts are in the model (public questions, 6% to 44%) and does nothing where they are not (private, 0% to 0%).
Chain of thought against the plain loop: acting brings the private facts within reach (0% to 67%), but the loop loses many public questions it should win, mostly by repeating itself (16 of 36 runs hit the budget).
The ablations: the two fixes repair different failures. The guard alone fixed every public question, because there the model already had what it needed and only had to stop. On private questions the guard mostly turned a stuck run into an honest "unknown", which is cheaper but not more correct. The improved tool alone fixed most private questions, because the model could finally find the depot entries, but runs still went in circles (9 hit the budget). Only both together fixed everything. Neither fix touches the model.
The repetition failure of the plain loop, and the two fixes that live in the loop and the tool rather than the model.
What function calling changed, and a complete modern loop#
When providers added structured tool calling to their APIs in 2023, the regular-expression parser disappeared. The model returns a tool call with a name and JSON arguments; your code runs it and sends back the result; the loop is otherwise the same. The hard parts moved from the text format to the loop (budgets, guards, stop conditions, error handling) and to the tools (what they return on a miss).
The script ch3_structured_loop.py is the loop written the way production loops are written today. It is one file of about 270 lines, comments included, runs against the real API, and has every piece this chapter talks about:
native tool calls: the model returns structured calls; several can arrive in one turn;
schema-checked arguments: each tool declares a JSON Schema, and the loop validates arguments before running anything;
an explicit state object: messages, step and token counts, calls already seen, and a status;
two budgets: model steps and total tokens; hitting either ends the run with a named status;
error handling: invalid JSON, an unknown tool, a schema violation or a tool exception becomes a tool result the model can read; API errors are retried with back-off;
a trace: one JSON line per event.
python
# code/agents/ch3_structured_loop.py (abridged): the loopdef run(question): s = RunState(question, messages=[{'role': 'system', 'content': SYSTEM}, {'role': 'user', 'content': question}]) while s.status == 'running': if s.steps >= MAX_STEPS: s.status = 'step_budget'; break if s.tokens_in + s.tokens_out >= MAX_TOKENS: s.status = 'token_budget'; break msg = call_model(s) # retries with back-off; None after three failures s.steps += 1 if msg is None: s.status = 'api_error'; break s.messages.append(assistant_message(msg)) for c in msg.tool_calls or []: result, args = execute(c.function.name, c.function.arguments) # parse, look up, validate, run; never raises key = (c.function.name, json.dumps(args, sort_keys=True)) if result['ok'] and key in s.seen: # the guard result['note'] = 'repeated call, result unchanged; answer with what you have' if result['ok'] and c.function.name == 'final_answer': s.answer, s.evidence, s.status = args['answer'], args['evidence'], 'answered' s.event('tool', tool=c.function.name, args=args, ok=result['ok']) s.messages.append({'role': 'tool', 'tool_call_id': c.id, 'content': json.dumps(result)}) return s
The final answer is itself a tool, final_answer, with a schema: a short answer string and a non-empty list of evidence entry names. That turns "the model stopped talking" into a typed event the loop can check.
On the depot question the loop looked up the company and the four depots one by one, computed 120 minus 41 and answered "The Rotterdam depot; it has 79 more trucks than the Leeds depot": seven model calls, 4,352 tokens, inside budgets of eight calls and 20,000 tokens. On the twelve questions, three runs each, it got 34 of 36 (94%), using 3.9 model calls and about 2,009 tokens per run on average. One of the two misses is worth remembering: on Everest against Kilimanjaro it looked up both heights (8,849 m and 5,895 m) and then called the calculator with 8848 - 5895, the height it remembered. The tool result was in the context and the model still used its prior. The other miss was the strict grader again ("Rotterdam (2011)" contains an extra number). Section 3.5 shows a check that catches the first kind.
Research notes: the ReAct paper in detail (optional reading)
Limitations the authors state. Prompting results were "still significantly far from domain-specific state-of-the-art approaches"; the smaller PaLM models prompted with ReAct did worst of the four methods, and only fine-tuning on 3,000 ReAct trajectories made them competitive; the repetition failure was attributed to greedy decoding and left open; and the study used three Wikipedia operations and two text games.
Rows changed: control across calls and context management.
The second Northwind question needs four look-ups that do not depend on each other. The structured loop above did them one per model call: seven calls, and each call re-read everything before it. A planner can write all four look-ups at once:
plain text
#E1 = lookup_entry["Rotterdam depot"]#E2 = lookup_entry["Gdansk depot"]#E3 = lookup_entry["Porto depot"]#E4 = lookup_entry["Leeds depot"]#E5 = solver: which of #E1..#E4 has the most trucks; calculate its trucks minus the trucks in #E4
An executor runs #E1 to #E4, in parallel if it likes, without calling the model at all; one last model call reads the plan and the four results and answers. That is two model calls instead of seven. (This plan is illustrative; we did not run a planner on it. The token arithmetic below is from ch3_cost.py.)
A loop re-reads a growing context on every call; plan-and-execute plans once, runs steps with small contexts and answers once.
Three 2023 papers built the idea up. Plan-and-Solve put "devise a plan, then carry it out" into one prompt. ReWOO ("Reasoning WithOut Observation") separated a planner, workers and a solver, with variables like #E2 so the planner never waits for a result. LLMCompiler dispatched the plan's steps in parallel as their inputs became ready, reporting latency speed-ups of up to 3.7 times over ReAct.
ch3_cost.py counts tokens for a task of n tool steps under stated planning numbers (a 1,500-token prompt, 150-token thoughts, 300-token observations), with no API calls. A loop that sends the full history every time sends roughly
n⋅P+2n(n−1)(T+O)
input tokens, where P is the prompt, T a thought and O an observation. The second term is the re-reading, and it grows with the square of the number of steps. A plan-and-execute design sends P to the planner, a small fixed brief to each executor step, and P plus the evidence to the solver: linear in n.
This comparison rests on two assumptions, and production systems often change both:
Full replay. The quadratic term assumes every call resends the whole growing history. Compaction (keeping only the last few steps in full, or a summary) makes the loop roughly linear, at the price of forgetting detail.
Full price for repeated tokens. Many providers offer prompt caching: a repeated prefix is billed at a fraction of the normal input price. That changes the bill a lot, but not the size of the context the model must attend to.
Billed-equivalent input tokens against tool steps under the planning numbers of ch3_cost.py: full replay, compaction, caching (at an assumed 10% price) and plan-and-execute.
At five steps the full-replay loop sends 15,750 input tokens and plan-and-execute 8,150; at twenty steps the ratio is 5.5 to one (126,000 against 22,850). With caching at our assumed 10% rate the twenty-step loop costs the equivalent of 22,050 tokens, about the same as the plan; with compaction it sends 57,150. So "ReAct is quadratic" is a statement about full replay at full price, not about ReAct. What a plan still buys under caching is fewer sequential model calls (latency) and an artefact a person can review before anything runs.
Loop (ReAct)
Plan-and-execute
Good for
few steps, each depending on the last; exploratory tasks
known task shapes with several independent steps; plans a person should approve
Costs
a model call per step; context growth under full replay
a planner that must know the tools; replanning adds cost back
Fails by
repetition, drift, budget exhaustion
a wrong plan executed in full before anyone notices
Measure
steps per task, repeat rate, tokens per task
dependency width (the parallel speed-up available), plan error rate, replanning rate
Research notes: Plan-and-Solve, ReWOO and LLMCompiler in detail (optional reading)
3.5 Checking the answer: who tells the agent it was wrong?#
Row changed: verification.
Go back to the structured loop's miss: it looked up 8,849 m and calculated with 8,848. Could the agent itself have caught that? Two kinds of check are possible.
Ask the model to review its answer. It sees the same context and the same prior that produced 8,848. Nothing new enters the loop.
Check the answer against the evidence. A few lines of code can test that every number fed to the calculator appears in an earlier observation. 8,848 appears in no observation, so the check fails and the loop can say exactly what is wrong: "8,848 is not in any looked-up entry; Mount Everest's entry says 8,849."
The second check brings in information the model did not use. That difference is the whole of this section.
The reflection loop. Whether it helps depends on the check: the model judging itself, or an external signal such as tests, a compiler, a search or a person.
Reflexion (Noah Shinn and colleagues, March 2023) let an agent retry after a failure with a short written reflection on what went wrong. On programming tasks the evaluator was a test suite the model wrote itself, run by a machine.
Reflexion's headline (GPT-4 on HumanEval, 80.1% to 91.0%) came with one honest loss: on MBPP the score fell from 80.1% to 77.1%, which the authors attribute to self-written tests that passed wrong code. Six months later Jie Huang and colleagues at Google DeepMind re-examined the self-correction results of 2023 and pointed out that several used an oracle: the true label decided when to stop retrying. Without it, asking a model to critique and revise its own reasoning made accuracy fall.
The two results fit together: Reflexion's gains came where a machine ran tests; Huang et al.'s losses came where the model had nothing but itself. Self-Refine (Madaan et al., 2023) found the same split from the other side: large gains on style and coverage tasks a model can see in its own text (dialogue, sentiment rewriting) and almost none on maths. CRITIC (Gou et al., 2023) supplied the constructive version: let the critic call tools (a search engine, a Python interpreter), and the gains come back; remove the tools and they mostly disappear. The research notes below give the tables.
The careful conclusion is conditional. Self-feedback can improve some outputs, especially where the defect is visible in the text, but it does not establish correctness. When correctness matters, use a check that is grounded in something independent of the model's own judgement.
Our experiment: repair with and without an external check#
ch3_reflexion.py gives gpt-5-mini ten small coding tasks with precise specifications (an expression evaluator, a CSV line parser, a cache with expiry, a cron matcher, a slug generator, an edit distance, numbers to British English words, a topological sort, a strict Roman numeral parser, a semantic version parser). Each task is attempted five times, so there are 50 first attempts.
Each task has two test suites:
a feedback suite, the project's own tests, which a repair loop may run, show to the model and use to decide when to stop;
a held-out suite: different inputs for the same rules in the specification, never shown to the model and never used to stop anything. Every number below is measured on the held-out suite.
Both suites were first checked against reference solutions written for the script (all 20 pass), so a failure means the code is wrong, not the test.
Then every first attempt, right or wrong, goes through two repair strategies, so we can count fixes and breakages:
self-review: two rounds of "review your function against the specification; if you find a mistake, reply with the corrected code, otherwise reply with the same code unchanged". No test is run or shown.
feedback-test loop: run the feedback suite; if it fails, show the first failing test and ask for a fix; up to two more attempts; stop as soon as the feedback suite passes.
Fifty first attempts graded on held-out tests: self-review left the pass rate at 86% (one fix, one breakage); the feedback-test loop raised it to 96% (five fixes, none broken).
Held-out pass rate
Wrong to right
Right to wrong
First attempt
43/50 (86%)
After two rounds of self-review
43/50 (86%)
1
1
After the feedback-test loop
48/50 (96%)
5
0
What happened, concretely:
Self-review was busy but not useful. It changed the code in 45 of 50 attempts (38 already right), yet the pass rate did not move: it fixed one cron matcher and broke one correct cache.
The feedback loop only touched what its check rejected. Six attempts failed the feedback suite. Four of them were caches whose get restarted the expiry time, which the specification forbids (only set restarts it). The failing test read the entry at 9.9 seconds and expected it gone at 10; with that in hand the model fixed three of the four. The other fixes were an expression evaluator (unary minus with powers) and a Roman numeral parser that accepted "XCX".
The check was not perfect. The feedback suite accepted 44 first attempts, and one of those failed the held-out tests (a cron matcher that rejected the range "5-7" in the day-of-week field). The loop never looked at it again, so it stayed wrong. A check's false accepts cap any loop built on it.
If each attempt passes with probability p and a verifier can tell right from wrong, the chance that at least one of k attempts passes is
P(success within k)=1−(1−p)k
With p=0.5 and k=3 that is 88%. The expected number of attempts is p1−(1−p)k, and dividing expected cost by the success probability gives a cost per correct answer of c/p for any k.
This formula assumes three things, and real loops often break all three:
Independent attempts. A feedback loop deliberately makes attempt 2 depend on attempt 1; that is the point of showing the failure.
Constant success probability and cost. The second attempt's chance differs from the first's (better with good feedback, sometimes worse), and its context, so its cost, is larger.
A perfect verifier. Real checks accept some wrong answers and reject some right ones, as our feedback suite did.
Treat the formula as a rough guide for independent resampling with a reliable check, and measure the real curve.
r = the probability the verifier accepts a correct candidate,
f = the probability the verifier accepts a wrong candidate (the false-accept rate).
The probability that an accepted answer is correct follows from Bayes' rule:
P(correct∣accepted)=pr+(1−p)fpr
Worked example: p=0.9, r=1, f=0.2 gives 0.9+0.1×0.20.9=0.920.9≈97.8%. A false-accept rate of 20% does not cap precision at 80%: precision depends on how many wrong candidates reach the verifier in the first place. The same verifier with p=0.3 gives 0.3+0.7×0.20.3≈68.2%. An acceptance means less on hard tasks, where most candidates are wrong.
P(correct | accepted) against the share of correct candidates, for two verifiers and none. The worked example, p = 90% and f = 0.2, gives 97.8%.Research notes: Reflexion, Self-Refine, CRITIC and the self-correction debate (optional reading)
The paper's other results follow the same rule. On ALFWorld the evaluator was a heuristic that detected repeated actions and over-long trajectories, external to the model, and success rose by about 22 points over twelve trials. On HotpotQA the evaluator was exact match against the gold answer, which is an oracle; Huang et al. later argued that most of the gain on such tasks comes from the oracle rather than the reflection.
Probability of success within k attempts for per-attempt pass rates of 30, 50 and 70 percent, under the formula's assumptions: independent attempts, the same pass rate on every attempt and a perfect verifier.
3.6 Choosing among candidates: best-of-n and search#
Row changed: verification and selection.
Without tools, the three direct answers to the Porto and Leeds question were "6 years", "9 years" and "5": a vote would only pick the most popular invention. On Everest, chain of thought gave 2,953 three times: a vote would be unanimous and wrong. Choosing among candidates helps only when the candidates contain the right answer and something can recognise it.
With the tools, the picture changes. Run the structured loop five times on the depot question and keep an answer only if a check passes (the evidence entries exist, every number used appears in an observation, the arithmetic recomputes). Now the selection has a real signal.
The idea predates the agent literature. Cobbe and colleagues at OpenAI trained a model to judge solutions to grade-school maths problems and used it to pick among many samples.
Best-of-n has a subtle limit, shown in the second table of the verifier screenshot above. If you return the first accepted candidate, more candidates raise the chance of returning something, but under independent candidates the precision of what you return stays at pr+(1−p)fpr, whatever n is. Two things make it worse in practice: correlated errors (the same wrong answer in many samples, which a lenient verifier then sees many times), and picking the top-scored candidate when the scorer has blind spots. Section 3.8 adds a further finding: how much extra sampling helps depends on how hard the question is.
Tree of Thoughts (Shunyu Yao and colleagues, May 2023) searched over partial solutions instead of complete ones: propose several next steps, score each ("sure", "maybe", "impossible"), expand only the promising ones.
Tree of thoughts: propose several next steps, score them, expand the promising ones; with the call arithmetic of ch3_cost.py and the paper's Game of 24 numbers.
The task explains the gain: a small search space, cheap and fairly reliable scoring of partial states, and a first move that decides everything. The paper's own accounting puts the tree at about $0.74 per puzzle against $0.47 for a hundred chains (2023 GPT-4 prices), so per solved puzzle the two cost about the same ($1.00 against $0.96). On a task a single chain solves 90% of the time, the same tree would mostly be waste. LATS (Zhou et al., 2023) extended the search to actions and environment feedback; its results, read in the notes, came with an oracle setup and around 67 model calls per question.
Cost and latency of seven patterns relative to a direct answer, under the planning numbers of ch3_cost.py. Divide by accuracy before deciding.
What to measure before choosing among candidates: the single-attempt pass rate (near 0 or near 1, selection cannot help), the verifier's false-accept rate, and, for a tree, the step at which failed chains go wrong and the nodes expanded per solved task.
Research notes: Tree of Thoughts and LATS in detail (optional reading)
3.7 Changing the action: code, interfaces and fixed pipelines#
Rows changed: action representation, and control across calls.
The depot question took the structured loop four look-up turns. If the action could be a short program, one turn would do:
python
trucks = {d: lookup_entry(d)['trucks'] for d in ['Rotterdam depot', 'Gdansk depot', 'Porto depot', 'Leeds depot']}best = max(trucks, key=trucks.get)print(best, trucks[best] - trucks['Leeds depot']) # Rotterdam depot 79
The model writes this once; a sandbox runs it and returns one line. (This is an illustration; the chapter's scripts do not run model-written code.) That is the idea behind code as action.
Code as action against JSON tool calls, and the Agentless counterpoint: a fixed pipeline of localisation, repair and validation.
Two other results from 2024 belong next to CodeAct, because together they say what to change first.
The interface matters. SWE-agent (John Yang and colleagues at Princeton) gave a model a purpose-built set of commands for working in a code repository: a file viewer that shows a window of lines, a search that summarises its matches, an edit command that checks syntax. Changing one interface element at a time moved the resolved rate by several points with the model held fixed.
In its ablations a search command that summarises matches was worth six points of resolved rate over one that lists them all, and showing a 100-line window instead of whole files was worth about five. The same paper's behavioural analysis produced a phrase worth keeping: agents succeed quickly and fail slowly. Successful runs finished at a median of 12 steps and $1.21; unsuccessful ones averaged 21 steps and $2.52. A long run is more likely lost than about to succeed, which is an argument for setting the step budget from successful runs.
Sometimes no agent is needed. Agentless (Chunqiu Steven Xia and colleagues, July 2024) fixed GitHub issues with a fixed three-phase pipeline (find the location, sample many patches, validate with tests) in which the model never chooses the next step.
Read the three together. CodeAct changed the action format; SWE-agent changed what the tools return; Agentless removed model-chosen control and kept sampling plus verification. None says "agents are good" or "agents are bad". When the steps of a task are known, a fixed pipeline with a check can be cheaper and more reliable; when the model must choose, give it a good interface and, where composition matters, the option to write code. Six months after Agentless, Anthropic reported 49% on SWE-bench Verified with a minimal scaffold in which the model chose every step, with a stronger model. The right amount of fixed structure appears to move with model capability and with how verifiable the task is.
Research notes: CodeAct, SWE-agent, Agentless and the SWE-bench scaffolds in detail (optional reading)
Most papers above used 2022 and 2023 models. Three later results qualify their conclusions. None of them overturns the practical advice; each changes how strongly it should be stated.
Self-correction can be trained. Huang et al. showed that prompted models of 2023 got worse when asked to correct themselves. Aviral Kumar and colleagues at Google DeepMind (SCoRe, September 2024) trained a model with multi-turn reinforcement learning on its own correction attempts. Their table measures exactly the two quantities of our Section 3.5 experiment: wrong-to-right (Δi→c) and right-to-wrong (Δc→i).
The best use of extra compute depends on difficulty. Charlie Snell and colleagues (August 2024) compared ways of spending test-time compute: sampling many answers and searching with a verifier against letting the model revise its answer in sequence.
Reasoning models' chains are more faithful, but far from fully. Yanda Chen and colleagues at Anthropic (2025) repeated the Turpin-style test on reasoning models: insert a hint into the prompt, find cases where the hint changed the answer, and check whether the chain of thought mentions the hint.
The papers measured benchmarks. Engineering write-ups describe systems with users. This section keeps two things apart: what the company reports (its own numbers, quoted, attributed to the exact system they describe) and our reading of which patterns the system uses. Where a page gives no number, none is given here. The full excerpts are in the research notes at the end of the section.
The patterns companies describe in their own write-ups, as we read them. Klarna is left out because its architecture was not published.
Company, system
What the company reports
Our reading of the pattern
Uber, Genie (security and privacy channels)
an LLM judge against a golden set cut evaluation from weeks to minutes; a relative 27% more acceptable answers and 60% less incorrect advice
a fixed retrieval pipeline; progress came from cheap evaluation, not a new loop
Amazon, Q Developer Java upgrades
tens of thousands of applications moved to Java 17; developers review the plan before it runs
plan-and-execute with a human approving the plan
Spotify, Fleet Management and Honk
Fleet Management automates about half of Spotify's pull requests since mid-2024; Honk, the coding agent inside it, has produced 1,500+ merged pull requests; a judge vetoes about a quarter of sessions
a home-made loop replaced by a coding agent with few tools, a verifiable goal and a judge
DoorDash, Dasher support
a two-tier guardrail cut hallucinations by 90% and severe compliance issues by 99%; a stronger guardrail was too slow and costly
a fixed pipeline with an external critic and a human fallback
LinkedIn, SQL Bot
validators check tables exist and run EXPLAIN; errors go to a correction step with tools
reflection on an external signal (the database)
Stripe, Minions
at most two rounds of CI per run, because CI costs tokens, compute and time
a state machine of fixed steps and agent steps; tests as the critic, capped
Airbnb, test migration
75% of 3,500 files migrated in four hours, 97% after four days of tuning, the rest by hand
a fixed per-file pipeline with validation errors fed back and capped retries
Spotify's numbers are easy to misattribute. "Around half of Spotify's pull requests" describes Fleet Management, the system that applies automated changes across thousands of repositories and existed before any model was involved. Honk, the coding agent added inside it, accounts for the title's "1,500+" merged AI-generated pull requests. The team reports that its first home-made loop "tended to get lost when it filled up its context window", and that its LLM judge, which vetoes about a quarter of sessions, has not yet been evaluated. Our reading: context growth and a missing check broke the first loop, and a judge is a check of unknown precision until it is measured.
Klarna: cost savings did not settle the quality question#
In February 2024 Klarna reported that its AI assistant handled two-thirds of customer service chats in its first month, "the equivalent work of 700 full-time agents". In May 2025, as Fortune reported from a Bloomberg interview, its chief executive said that "cost unfortunately seems to have been a too predominant evaluation factor" and that the result was "lower quality", and Klarna began hiring people for customer service again. The architecture was never published, so nothing here says which pattern it used. What the story shows is about measurement: reporting cost and volume first left the quality question open, and it had to be answered later.
Research notes: the company write-ups in detail (optional reading)
Uber. Three systems cover most of this chapter's patterns.
Amazon.
Spotify.
DoorDash.
DoorDash's later posts trace the same team moving from fixed workflows to single agents and then to several agents. A November 2025 post names the single agent's limit as "context pollution", and a June 2026 post on the consumer DoorDash Assistant reports that "the largest potential production-failure category is grounding" and that "the fix in each case has been to route the agent's claim through a tool call against the system of record".
LinkedIn.
Stripe.
Airbnb. The March 2025 post "Accelerating Large-Scale Test Migration with LLMs" (the Airbnb Tech Blog blocks automated screenshots, so it is quoted) describes migrating "nearly 3.5K React component test files" between testing frameworks, a job estimated at "1.5 years of engineering time" and finished "in just 6 weeks". Each file went through a fixed pipeline ("we modeled this flow like a state machine") that would "retry steps multiple times until they passed or we reached a limit", with "the validation errors and the most recent version of the file" in each retry prompt. The team reports "75% of our target files in just four hours", "from 75% to 97%" after four days of tuning, and the "remaining 3%" finished by hand. Our reading: a fixed pipeline, reflection on an external signal, capped retries and a human tail.
Back to Northwind. Suppose the company wants an assistant that answers staff questions from its encyclopaedia. Here is how this chapter's measurements would guide the build, one decision at a time.
A labelled set and the simplest system. Twelve typed questions, an exact grader. Direct answering: 3%, failing on missing facts.
Missing facts call for tools in a loop. Plain ReAct: 56%; the traces show repetition and tool misses.
Fix the loop and the tool first. Guard plus improved tool: 100%. A structured loop: 94%, with misses an evidence check would catch.
Add the check the misses call for: an evidence check costs a few lines and no model calls.
A plan only if steps are many and independent, priced with caching and compaction included.
Best-of-n or search only if single attempts are middling and a reliable check exists.
That sequence is a starting heuristic, not a law. The five dimensions of Section 3.1 are independent, so you can skip or combine steps when the task makes the need obvious: code with a test suite can start at "reflection on tests"; a puzzle with cheap state scoring can start at search; a task whose steps are fully known can start, and stay, as a fixed pipeline.
A starting heuristic, not a fixed sequence: add one thing per measured failure, skip or combine steps when the task makes the need obvious, and re-measure.
Task property
How to measure it
What it argues for
Verifiability: can an answer be checked without knowing it?
list the available checks and measure their false-accept rates
with a strong check: reflection on it, and best-of-n; without one: short loops and a person at the end
Step count
the distribution of tool calls in solved runs
one to three: a direct call or a short loop; more, with independent steps: a plan
Knowledge
the direct-answer baseline
high: chain of thought may be enough; low: tools, and check that tools do not hurt the easy cases
Latency budget
the product's p95 against the pattern's sequential depth
tight: parallel candidates or a plan with parallel steps
Cost of a wrong answer
what happens downstream
high: spend on verification; low: the cheapest pattern that meets the bar
Shape of failures
the step at which failed runs went wrong
repetition: a guard; tool misses: better observations; early fatal choices: search; late, checkable errors: reflection
ch3_cost.py puts the cost per correct answer of seven patterns side by side under planning numbers:
Read the second table with care. The three accuracies marked "mea" are measured (Section 3.3, the plain loop); the rest are assumptions to show the arithmetic. Under those numbers, at frontier-model prices, chain of thought costs about 4.4 cents per correct answer and the plain loop about 10.9 cents, because almost half its runs were wasted; a plan assumed to reach 80% would cost about 4.6 cents and a tree assumed to reach 90% about 28 cents. Put in your own accuracies and the ranking can change; the formula will not.
The one rule of the chapter, stated as a condition: add a pattern when a measured failure calls for it, and keep it only if accuracy, cost per correct answer and latency still justify it after you re-measure.
Your wiki agent uses a ReAct loop with a ten-step budget, and 30% of failed runs hit the budget. Design three measurements that separate repetition, unhelpful tool misses and genuinely long tasks, and name the fix each one points to.
A team wants to add "the model reviews its answer before replying" to a customer-facing agent. Using Huang et al., SCoRe's Table 2 and our repair experiment, write the one-paragraph case for measuring wrong-to-right and right-to-wrong changes first, and list three external checks available in a typical support system.
Re-run ch3_cost.py with your own prompt size, observation size, prices and cache discount. At what step count does plan-and-execute become cheaper than the loop with full replay, and does it still win with caching?
Your verifier accepts 5% of wrong answers and all right ones. Compute P(correct∣accepted) for p=0.8 and p=0.2. Which of your task types sits nearer each value, and what does that imply about where to spend on a better verifier?
Take the Spotify and Airbnb cases. For each, write one sentence on what the company reports and one on your reading of the pattern, and name one number you would want that neither post gives.
A reasoning pattern is a choice on one or more of five independent dimensions: reasoning inside one call, control across calls, action format, verification and selection, and context management. They combine; they are not a ladder.
Chain of thought helps a model use what it knows (public questions: 6% to 44% in our run) and cannot add facts it lacks (private: 0% to 0%). Self-consistency returns the plurality answer; it fails when errors are correlated.
ReAct grounds the model in tool results but fails by repetition and unhelpful tool misses. In our ablation the plain loop scored 56%, the guard alone 75%, the improved tool alone 75%, and both 100%. Its "0% hallucination" is a share of sampled failed runs in one study, not an overall rate.
A modern structured loop (schema-checked tool calls, explicit state, budgets, typed final answer, trace) scored 34 of 36 on the same questions; its misses were a prior overriding an observation and a strict grader.
Plan-and-execute cuts sequential calls and, under full replay, tokens; caching and compaction change the token comparison, so state the assumptions.
Reflection is worth what its check is worth. In our held-out experiment, self-review left 43 of 50 unchanged in total (one fix, one breakage); the feedback-test loop reached 48 of 50 (five fixes, none broken).
A verifier's false-accept rate is not its precision: p=0.9, r=1, f=0.2 gives about 97.8% precision, but p=0.3 gives 68.2%. More candidates raise the chance of returning something, not its precision.
Newer work qualifies the 2023 results: self-correction can be trained (SCoRe), the best use of test-time compute depends on difficulty (Snell et al.), and reasoning models' chains are more faithful but often still omit what drove the answer (Chen et al.).
In production write-ups, keep a company's reported numbers apart from your reading of its pattern.
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, John Schulman. Training Verifiers to Solve Math Word Problems. arXiv October 2021.
Sehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee, Michael W. Mahoney, Kurt Keutzer, Amir Gholami. An LLM Compiler for Parallel Function Calling. ICML 2024 (arXiv December 2023).
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal Behbahani, Aleksandra Faust. Training Language Models to Self-Correct via Reinforcement Learning. arXiv September 2024.
code/agents/ch3_react.py: twelve two-hop questions, six system configurations (including the four-way ablation of the guard and the improved tool), three repeats, a typed exact grader with a self-test; results in results/ch3_react.json and the two ch3_react*_stdout.txt logs.
code/agents/ch3_structured_loop.py: a complete structured-tool loop (native tool calls, schema validation, explicit state, step and token budgets, error handling, a JSON-lines trace); results in results/ch3_structured_loop.json and results/ch3_structured_trace.jsonl.
code/agents/ch3_reflexion.py: ten coding tasks with feedback and held-out test suites (both checked against reference solutions), self-review against a feedback-test loop on every first attempt; results in results/ch3_reflexion.json.
code/agents/ch3_cost.py: the cost and latency arithmetic, cost per correct answer, the retry formula and its assumptions, token growth with caching and compaction, and the verifier precision and best-of-n tables; results in results/ch3_cost.json.
This chapter kept one agent and changed how it thinks, acts and checks. Chapter 4 keeps the thinking and changes the plumbing around it: workflow patterns in which code, not the model, decides the path. Prompt chaining, routing, parallel fan-out, orchestrator and workers, and evaluator and optimiser are ways of arranging several model calls so that each is small, checkable and cheap. Several systems in this chapter (Uber's review pipeline, Stripe's state machine, Airbnb's migration) are workflows with an agent inside one step, and Snell et al.'s finding that compute should follow difficulty is, in practice, a routing decision.