Agentic Systems · Part 2 · Patterns

Chapter 3 · Reasoning patterns: how one agent thinks, acts and corrects itself

Reasoning patterns for one agent: chain of thought, ReAct, plan-and-execute, reflection with an external check, tree search and code as action.

Goal: by the end of this chapter you can take one task, try it with several loop shapes (think, act, plan, check, choose among candidates), say what each shape changes, what it costs and how it fails, and decide with a measurement which one to keep. You will read the key papers behind each pattern, see three small experiments run against a real model, and have one complete modern tool loop you can copy.


3.1 One task, many loop shapes

Chapter 1 drew the agent as a loop and Chapter 2 took the loop apart into blocks. This chapter keeps the blocks and changes one thing: the shape of the loop. How many times is the model called? Does it write its reasoning down? Does it look things up as it goes, or plan everything first? Does anything check the answer? Does it try several answers and pick one?

We will follow one example the whole way through.

Two questions come back again and again:

  1. How many years after the Porto depot opened did the Leeds depot open? The answer is 3 (2020 minus 2017). It needs two look-ups and a subtraction.
  2. Which depot has the most trucks, and how many more does it have than the Leeds depot? The answer is Rotterdam, with 79 more (120 minus 41). It needs four independent look-ups, a comparison and a subtraction.

Here is what each shape does with the first question, before any theory:

ShapeWhat the system does with the Porto and Leeds questionWhat happened in our runs (gpt-5-mini)
Direct answerone call, answer only"6 years", "9 years", "5": invented every time
Chain of thoughtone call, write the steps firsttwo of three runs said "I don't have the opening years"; one said "2 years"
ReAct loopthink, look up, observe, repeatright in one run of three; two runs went round in circles until the step budget ran out
Loop + guard + better toolthe same loop, with two engineering fixesright in three runs of three
Plan firstwrite all look-ups, run them, answer oncenot run on this question; Section 3.4 counts its cost
Check the answerrecompute from the evidence before replyingcatches a wrong number before the user sees it (Section 3.5)

The rest of the chapter explains each row: the idea, the paper that introduced it, what it costs, how it fails and what to measure.

Six shapes of the loop, same model and same toolsthink onlychain of thoughtqthoughtsansact onlytool calls, no thoughtscallcallcallansinterleaveReAct: think, act, observethinkactobserverepeat until doneplan firstplan, execute, solveplanstep 1step 2step 3solvereflecttry, check, critique, retryattemptcheckcritiqueretry with the critiquesearchbranch, score, prunescore,keep bestthe shapes are not a sequence: a real agent combines them (a plan over a loop, a check after it).Each one costs calls, tokens and seconds; this chapter asks when each one pays for itself.
Six shapes of the loop with the same model and tools: think only, act only, interleave (ReAct), plan first, reflect, and search.

The shapes are choices on five separate dimensions

It is tempting to read the six shapes as a ladder, each one "more advanced" than the last. That is the wrong picture. Each pattern changes one or two separate design choices, and a real agent picks an option on each.

Five design dimensions: each agent picks one or more options per rowreasoning in one calldirect answerchain of thoughtreasoning modelvote over samplescontrol across callsone callfixed pipelineloop (ReAct)plan, executeaction formattext, parsedJSON tool callcode in sandboxno toolsverificationnoneself-reviewexternal checkbest-of-n, searchcontextfull replaycompactionbounded per stepprompt cachinghighlighted: the structured loop of Section 3.3 (reasoning model, loop, JSON tool calls,a typed final check, full replay). CodeAct changes only row 3; plan-and-execute rows 2 and 5;best-of-n row 4. The rows are independent, so the patterns of this chapter combine.
Five independent design dimensions of a single agent. A pattern is a choice in one or two rows; the highlighted cells are the structured loop of Section 3.3.
DimensionThe question it answersPatterns that change it
Reasoning inside one callhow does the model think before it answers?chain of thought, self-consistency, reasoning models
Control across callswho decides the next step?ReAct, plan-and-execute, Agentless
Action representationwhat does an action look like?function calling, CodeAct
Verification and selectionwhat decides that an answer is good?reflection, CRITIC, best-of-n, tree search
Context managementwhat does each call see?ReWOO, compaction, prompt caching

So "we use ReAct" says something about control and nothing about verification. Every section below says which row it changes.

Three numbers decide

Every pattern is judged on three numbers measured on a labelled set: accuracy, cost per task (tokens and tool calls) and latency (time to answer). The number that combines the first two is:

Most papers below used models that are now several generations old. Read their tables for the shape of each result (which pattern helps on which task, at what cost) and re-measure on your own model.

3.2 Thinking inside one call: chain of thought, voting and reasoning models

Row changed: reasoning inside one call.

Start with the public question "How many years passed between the completion of the Eiffel Tower and the opening of the Sydney Opera House?" (1973 minus 1889, so 84). Asked for the answer only, gpt-5-mini said 66, 72 and 75 in three runs. Asked to "think step by step" first, it wrote down 1889 and 1973, subtracted, and said 84 all three times. Nothing else changed: the same model, the same question, one call each.

Now the Northwind question about Porto and Leeds. Thinking step by step did not help. In two runs the model said, correctly, that it did not have the opening years; in the third it wrote "2 years". On the question about the chief executive it invented three different names in three runs. A scratchpad helps a model use what it knows. It cannot create a fact the model never saw.

standard promptingchain of thoughtquestionanswer"The answer is 27."one call, about 30 output tokens;the arithmetic has to happen insidethe model, with nothing written downquestionreasoning"23 - 20 = 3; 3 + 6 = 9"answerone call, hundreds of output tokens;each intermediate result is writtenand read back before the next stepself-consistency: sample several chains, return the most frequent answerquestionchain 19chain 29chain 327chain 49chain 512vote9 (3 of 5)the vote is a plurality: 9 wins with 3 of 5, and would still win with 2 of 5 if the wrong answersdisagree. Five chains cost five times the tokens of one. It helps when answers can be comparedexactly; it cannot help when the chains share the same mistake (correlated errors).
Standard prompting against chain of thought, and self-consistency: several sampled chains and a plurality vote over their final answers.

The paper

In January 2022 Jason Wei and colleagues at Google Research showed that a large enough model, given a few examples whose answers contain the working, writes working for new problems too and gets many more of them right. No training was involved.

On GSM8K (grade-school maths word problems) PaLM-540B went from about 18% to about 57%. The gain appeared only in large models; small models often got worse. Four months later Takeshi Kojima and colleagues showed that one sentence, "Let's think step by step", gets most of the gain with no examples.

Voting over several chains

If one chain is good, are several better? Xuezhi Wang and colleagues (2022) sampled many chains at a non-zero temperature, read off each final answer and returned the most common one. They called it self-consistency. With 40 samples, PaLM-540B on GSM8K rose from 56.5% to 74.4%.

Two details matter in practice.

Plurality, not majority. Suppose each chain is right 40% of the time, and wrong chains split evenly between two different wrong answers, 30% each. No answer has a majority, yet over many samples the right answer is the most frequent and wins. Voting needs the right answer to be more common than any single wrong answer, not more common than all wrong answers together.

Correlated errors. The argument above assumes the wrong chains scatter. If they make the same mistake, voting makes it more certain, not less. Our chain-of-thought runs show this: on Everest against Kilimanjaro all three chains answered 2,953 metres, because the model remembers Everest as 8,848 m while our encyclopaedia says 8,849 m. A vote over those three chains would return 2,953 with full agreement. Agreement measures consistency, not correctness.

Self-consistency was not the first result to spend extra computation at answer time. In October 2021 Karl Cobbe and colleagues at OpenAI generated many candidate solutions per problem and picked the one a trained verifier ranked highest (Section 3.6). Self-consistency replaced the trained verifier with agreement between samples.

Reasoning models: the scratchpad moved inside

From late 2024 the pattern moved into training. OpenAI's o1 series, then DeepSeek-R1 and others, were trained with reinforcement learning to produce a long chain of thought before answering, rewarded when the final answer could be checked.

prompted chain of thought (2022 to 2024)reasoning models (late 2024 onwards)who writes the reasoningthe prompt asks for it ("think step by step")where it appearsin the visible output, as ordinary tokenshow it was learnedimitation of human-written steps, or nonewhat you payoutput tokens you can see and countwhat you can inspectthe whole chain (faithful or not)what you controlformat, length, examples, when to thinkwho writes the reasoningthe model, trained by reinforcement learningwhere it appearshidden or summarised; the answer is shownhow it was learnedrewarded for verifiable answers (maths, code)what you payreasoning tokens billed as output, unseenwhat you can inspecta summary, if any; not the raw chainwhat you controlan effort setting; little elsewhat did not change: a loop still needs tools, a stop condition, a budget and a trace;a wrong answer is still wrong; and the counter-evidence later in this section (unfaithfulexplanations, weak self-correction) was measured on the visible kind of reasoning.
What reasoning models changed about chain of thought and what they did not. Visibility and controls differ across models and providers.

For an agent builder, much stays the same: the model still cannot look anything up, still needs a stop condition and a budget, and is still wrong some of the time. What changed is the economics (with some providers you pay for reasoning you cannot read in full) and the question ("should I ask it to think?" becomes "how much effort should I buy for this step?"). Our experiments use gpt-5-mini with its effort set to minimal, so the reasoning we study is the visible kind.

Is the chain the real reason?

This result was measured on 2023 models. Section 3.8 shows the 2025 follow-up on reasoning models: the chains became more faithful but were still far from fully faithful in the settings studied. The practical rule is conditional: a chain is useful for debugging, but where correctness matters, do not treat a plausible chain as evidence. Check the answer against something outside the text.

What to measure. Accuracy on the multi-step slice of your labelled set with and without a chain, output tokens per call, and, for a reasoning model, accuracy gained per dollar at each effort setting.

Research notes: the chain-of-thought papers in detail (optional reading)

These are the remaining excerpts and table readings behind Section 3.2. Nothing here is needed for the rest of the chapter.

3.3 Acting in a loop: ReAct

Row changed: control across calls (and, once tools appear, the action format).

Thinking cannot produce the Porto depot's opening year. Looking it up can. The simplest loop that looks things up is: think about what you need, call a tool, read the result, repeat. Here is what that loop did on the Porto and Leeds question when it worked:

plain text
Thought: I need the opening years of the Porto and Leeds depots of Northwind Freight.
Action: wiki[Northwind Freight]
Observation: Fictional logistics company founded in 2011 in Rotterdam. ... Related entries: Rotterdam depot, Gdansk depot, Porto depot, Leeds depot, ...
Action: wiki[Porto depot]      Observation: Northwind Freight depot opened 2017. Vehicles: 64 trucks. ...
Action: wiki[Leeds depot]      Observation: Northwind Freight depot opened 2020. Vehicles: 41 trucks. ...
Action: calc[2020-2017]        Observation: 3
Action: finish[3]

And here is the same plain loop on a public question, when it did not:

Terminal output of the trace log of ch3_react.py: the plain ReAct loop on the Eiffel Tower question looks up both entries, computes calc 1973 minus 1889 and observes 84, then computes it again, then starts the look-ups over until the seven-step budget runs out and the run is graded wrong; then the guarded loop with the improved tool on the Porto and Leeds question, which looks up the company, the two depots, computes 2020 minus 2017 and finishes with 3, graded correct; then the direct answer, 6 years, and a chain of thought that says it does not have the opening years

After step 3 the answer, 84, is in the context. The model computes it again, then starts the look-ups over, and the step budget runs out. This failure has a name in the paper that introduced the loop.

ReAct: thought, action, observation, repeatedThoughtActionObservationwriterun the toolreadthe model writes a thought and an action;the loop runs the tool and appends theobservation; exit with Action: finish[answer]the paper's text format (2022)Thought: I need the depot's opening year.Action: wiki[Porto depot]Observation: depot opened 2017. Trucks: 64.Thought: ... Action: finish[3]the same loop with function calling (now)assistant: tool_use wiki {name: "Porto depot"}user: tool_result "depot opened 2017 ..."the thought is optional text before the call,or hidden reasoning tokens, or absentthe loop shape is identical; what changed is who parses the action (a regular expression then,the API now) and whether the thought is shown at all.
The ReAct loop and its two surface forms: the 2022 text format parsed by a regular expression, and the same loop with structured tool calls.

The paper

Shunyu Yao (Princeton) and colleagues at Google posted ReAct in October 2022. Chain of thought could reason but could not look anything up, so it invented facts. Earlier acting agents could call tools but kept no reasoning between actions, so they lost track of the goal. ReAct put both in one loop. For question answering the tools were three operations over Wikipedia (search, lookup, finish) and the budget was seven steps.

Two readings matter. First, ReAct lost to chain of thought on HotpotQA (27.4 against 29.4); the best results combined the two with a fallback. Second, the paper's hand-labelled error study (Table 2, read in the research notes below) found the two failures our trace shows: "reasoning error", which the authors define to include failing to recover from repetitive steps (47% of the sampled ReAct failures), and non-informative searches that derail the run (23%). The same table is often quoted as "ReAct has 0% hallucination". That is too strong. The 0% is the hallucination category among sampled failed ReAct trajectories; among sampled successful ones, 6% were right with a hallucinated reasoning trace or facts. Grounding reduced hallucination in that study; it did not remove it.

Our experiment: six system configurations

The script ch3_react.py runs the twelve questions (six public, six about Northwind) through six configurations with gpt-5-mini at minimal reasoning effort, three times each, so every rate is over 36 runs. The four loop configurations isolate two engineering fixes:

  • Repeated-action guard (in the loop): if the model emits an action it has already run, the loop replaces the observation with "you have been round this loop; using only the observations above, finish now".
  • Improved tool (in the tool): an entry lists related entries by name (the company entry names its four depots as "Porto depot" and so on), several matches come back as entries instead of an error, and a miss explains how entries are named.

Grading is typed and exact. An answer is right only if it contains exactly the expected numbers, compared as numbers, and names every expected entity after explicit normalisation (accents removed, lower case, punctuation collapsed). An earlier draft of this script used substring matching, which marked "184" correct for 84. The script now starts with a self-test that shows the difference:

python
# code/agents/ch3_react.py (abridged): the typed grader and the guard
def grade(answer, expected):
    got = sorted(numbers_in(answer))                          # "2,954 m" -> [2954.0]; "three" -> [3.0]
    want = sorted(float(x) for x in expected.get('numbers', []))
    if len(got) != len(want) or any(abs(g - w) > 1e-9 for g, w in zip(got, want)):
        return False                                          # "184", "13", "It is not 84; the answer is 74." all fail
    text = f' {normalise(answer)} '
    return all(f' {normalise(e)} ' in text for e in expected.get('entities', []))

if (tool, arg) in seen:                                       # inside the loop
    repeats += 1
    if guard:
        obs = 'You have been round this loop once. Using ONLY the observations above, ... finish[answer] now.'

Terminal output of ch3_react.py: the grader self-test, where the old substring grader marks 184, 13 and It is not 84 the answer is 74 as right and the typed grader marks them wrong; a per-question table of correct runs out of three for six configurations; and the summary over 36 runs each: direct 3 percent, chain of thought 22, ReAct 56, ReAct plus guard only 75, plus improved tool only 75, plus both 100, with calls, tokens, cost, cost per correct answer, budget hits and repeat counts; total spend 0.1219 dollars over 715 calls

Six configurations, 12 questions x 3 repeats, gpt-5-mini (ch3_react.py)accuracy over 36 runs (typed exact grader)cost per correct answer, US centsdirect answerdirect answer: 3%3%chain of thoughtchain of thought: 22%22%ReAct, plain loopReAct, plain loop: 56%56%+ guard only+ guard only: 75%75%+ improved tool only+ improved tool only: 75%75%+ guard + improved tool+ guard + improved tool: 100%100%: 0.144c0.144c: 0.086c0.086c: 0.172c0.172c: 0.085c0.085c: 0.111c0.111c: 0.072c0.072cReAct, plain loop: public 44%, private 67%; 189 calls; 16 runs hit the step budget+ guard only: public 100%, private 50%; 134 calls; 0 runs hit the step budget+ improved tool only: public 56%, private 94%; 170 calls; 9 runs hit the step budget+ guard + improved tool: public 100%, private 100%; 150 calls; 0 runs hit the step budgetthe guard alone fixes the public questions (the model had the facts but kept repeating itself);the improved tool alone fixes most private ones (the model could not find the depot entries);both together fix all twelve. The conditions also differ in prompts, budgets and tools, so thiscompares system configurations, not loop shapes alone.
Six system configurations, 36 runs each: the guard and the improved tool each fix different questions; only both together answer all twelve.
ConfigurationAccuracy (36 runs)PublicPrivateCallsCost per correctRuns that hit the budget
Direct answer3%6%0%36$0.001440
Chain of thought22%44%0%36$0.000860
ReAct, original tool, no guard56%44%67%189$0.0017216
+ guard only75%100%50%134$0.000850
+ improved tool only75%56%94%170$0.001119
+ guard + improved tool100%100%100%150$0.000720

Read it in three steps.

  1. Direct against chain of thought: a scratchpad helps where the facts are in the model (public questions, 6% to 44%) and does nothing where they are not (private, 0% to 0%).
  2. Chain of thought against the plain loop: acting brings the private facts within reach (0% to 67%), but the loop loses many public questions it should win, mostly by repeating itself (16 of 36 runs hit the budget).
  3. The ablations: the two fixes repair different failures. The guard alone fixed every public question, because there the model already had what it needed and only had to stop. On private questions the guard mostly turned a stuck run into an honest "unknown", which is cheaper but not more correct. The improved tool alone fixed most private questions, because the model could finally find the depot entries, but runs still went in circles (9 hit the budget). Only both together fixed everything. Neither fix touches the model.
The most common ReAct failure in our run, and where the two fixes live1step 1wiki[Eiffel Tower]Obs: completed 18892step 2wiki[Sydney Opera House]Obs: opened 19733steps 3 and 4calc[1973-1889]Obs: 84calc[1973-1889] again4steps 5 to 7look-ups start overbudget of 7 runs outno answer: wrongthe answer (84) was in the context after step 3; the model kept acting instead of finishing(the "reasoning error" row of the ReAct paper, which includes failing to recover from repetitive steps)two fixes outside the model, measured separatelythe loop: a repeated-action guardthe loop sees the same (tool, argument) twice andreplaces the observation with "you have beenround this loop; using what you have, finish now"the tool: observations that name next stepsa company entry lists its depots as look-up names;several matches come back as entries, not errors;a miss explains how entries are namedplain loop 56%; guard only 75%; improved tool only 75%; both 100% (36 runs each)
The repetition failure of the plain loop, and the two fixes that live in the loop and the tool rather than the model.

What function calling changed, and a complete modern loop

When providers added structured tool calling to their APIs in 2023, the regular-expression parser disappeared. The model returns a tool call with a name and JSON arguments; your code runs it and sends back the result; the loop is otherwise the same. The hard parts moved from the text format to the loop (budgets, guards, stop conditions, error handling) and to the tools (what they return on a miss).

The script ch3_structured_loop.py is the loop written the way production loops are written today. It is one file of about 270 lines, comments included, runs against the real API, and has every piece this chapter talks about:

  • native tool calls: the model returns structured calls; several can arrive in one turn;
  • schema-checked arguments: each tool declares a JSON Schema, and the loop validates arguments before running anything;
  • an explicit state object: messages, step and token counts, calls already seen, and a status;
  • two budgets: model steps and total tokens; hitting either ends the run with a named status;
  • error handling: invalid JSON, an unknown tool, a schema violation or a tool exception becomes a tool result the model can read; API errors are retried with back-off;
  • a trace: one JSON line per event.
python
# code/agents/ch3_structured_loop.py (abridged): the loop
def run(question):
    s = RunState(question, messages=[{'role': 'system', 'content': SYSTEM}, {'role': 'user', 'content': question}])
    while s.status == 'running':
        if s.steps >= MAX_STEPS:                      s.status = 'step_budget'; break
        if s.tokens_in + s.tokens_out >= MAX_TOKENS:  s.status = 'token_budget'; break
        msg = call_model(s)                           # retries with back-off; None after three failures
        s.steps += 1
        if msg is None:                               s.status = 'api_error'; break
        s.messages.append(assistant_message(msg))
        for c in msg.tool_calls or []:
            result, args = execute(c.function.name, c.function.arguments)   # parse, look up, validate, run; never raises
            key = (c.function.name, json.dumps(args, sort_keys=True))
            if result['ok'] and key in s.seen:        # the guard
                result['note'] = 'repeated call, result unchanged; answer with what you have'
            if result['ok'] and c.function.name == 'final_answer':
                s.answer, s.evidence, s.status = args['answer'], args['evidence'], 'answered'
            s.event('tool', tool=c.function.name, args=args, ok=result['ok'])
            s.messages.append({'role': 'tool', 'tool_call_id': c.id, 'content': json.dumps(result)})
    return s

The final answer is itself a tool, final_answer, with a schema: a short answer string and a non-empty list of evidence entry names. That turns "the model stopped talking" into a typed event the loop can check.

Terminal output of ch3_structured_loop.py: six error paths exercised directly, each becoming a readable tool result (a name too short, an extra argument, an expression with words in it, division by zero, invalid JSON, an unknown tool); then the Northwind question about the depot with the most trucks, answered in seven model calls with look-ups of the company and four depots, a calculation of 120 minus 41, and a final answer of Rotterdam with 79 more trucks, 4,352 tokens, graded correct; then the same loop on the twelve questions three times, 34 of 36 correct

On the depot question the loop looked up the company and the four depots one by one, computed 120 minus 41 and answered "The Rotterdam depot; it has 79 more trucks than the Leeds depot": seven model calls, 4,352 tokens, inside budgets of eight calls and 20,000 tokens. On the twelve questions, three runs each, it got 34 of 36 (94%), using 3.9 model calls and about 2,009 tokens per run on average. One of the two misses is worth remembering: on Everest against Kilimanjaro it looked up both heights (8,849 m and 5,895 m) and then called the calculator with 8848 - 5895, the height it remembered. The tool result was in the context and the model still used its prior. The other miss was the strict grader again ("Rotterdam (2011)" contains an extra number). Section 3.5 shows a check that catches the first kind.

Research notes: the ReAct paper in detail (optional reading)

Limitations the authors state. Prompting results were "still significantly far from domain-specific state-of-the-art approaches"; the smaller PaLM models prompted with ReAct did worst of the four methods, and only fine-tuning on 3,000 ReAct trajectories made them competitive; the repetition failure was attributed to greedy decoding and left open; and the study used three Wikipedia operations and two text games.

3.4 Planning before acting

Rows changed: control across calls and context management.

The second Northwind question needs four look-ups that do not depend on each other. The structured loop above did them one per model call: seven calls, and each call re-read everything before it. A planner can write all four look-ups at once:

plain text
#E1 = lookup_entry["Rotterdam depot"]
#E2 = lookup_entry["Gdansk depot"]
#E3 = lookup_entry["Porto depot"]
#E4 = lookup_entry["Leeds depot"]
#E5 = solver: which of #E1..#E4 has the most trucks; calculate its trucks minus the trucks in #E4

An executor runs #E1 to #E4, in parallel if it likes, without calling the model at all; one last model call reads the plan and the four results and answers. That is two model calls instead of seven. (This plan is illustrative; we did not run a planner on it. The token arithmetic below is from ch3_cost.py.)

ReAct: one growing contextplan-and-execute: plan, then small contextscall 1toolcall 2 + 1 obstoolcall 3: prompt + 2 thoughts + 2 obstoolcall 4: prompt + 3 thoughts + 3 obsevery call re-reads everything before it:15,750 input tokens for 5 steps,126,000 for 20 (ch3_cost.py)plannerone big call#E1 = wiki[Porto depot]#E2 = wiki[Leeds depot]#E3 = calc[#E2.year - #E1.year]worker 1worker 2worker 3solverplan once, run the steps with small contexts(in parallel where the plan allows), answer once:8,150 input tokens for 5 steps, 22,850 for 20.the price: a wrong plan is noticed only at the end
A loop re-reads a growing context on every call; plan-and-execute plans once, runs steps with small contexts and answers once.

The papers, briefly

Three 2023 papers built the idea up. Plan-and-Solve put "devise a plan, then carry it out" into one prompt. ReWOO ("Reasoning WithOut Observation") separated a planner, workers and a solver, with variables like #E2 so the planner never waits for a result. LLMCompiler dispatched the plan's steps in parallel as their inputs became ready, reporting latency speed-ups of up to 3.7 times over ReAct.

The cost arithmetic, and its assumptions

ch3_cost.py counts tokens for a task of n tool steps under stated planning numbers (a 1,500-token prompt, 150-token thoughts, 300-token observations), with no API calls. A loop that sends the full history every time sends roughly

n⋅P+n(n−1)2(T+O)n \cdot P + \frac{n(n-1)}{2}(T + O)

input tokens, where PP is the prompt, TT a thought and OO an observation. The second term is the re-reading, and it grows with the square of the number of steps. A plan-and-execute design sends PP to the planner, a small fixed brief to each executor step, and PP plus the evidence to the solver: linear in n.

This comparison rests on two assumptions, and production systems often change both:

  • Full replay. The quadratic term assumes every call resends the whole growing history. Compaction (keeping only the last few steps in full, or a summary) makes the loop roughly linear, at the price of forgetting detail.
  • Full price for repeated tokens. Many providers offer prompt caching: a repeated prefix is billed at a fraction of the normal input price. That changes the bill a lot, but not the size of the context the model must attend to.

Terminal output of ch3_cost.py part four: ReAct input tokens against step count with full replay, with prompt caching at 10 percent of the input price, with compaction keeping three steps, and plan-and-execute: 2 steps 5,850, 2,745, 5,850 and 5,210; 5 steps 15,750, 4,950, 14,400 and 8,150; 10 steps 41,250, 9,525, 28,650 and 13,050; 20 steps 126,000, 22,050, 57,150 and 22,850

025,00050,00075,000100,000125,000251020ReAct, full replay, 2: 5,850ReAct, full replay, 5: 15,750ReAct, full replay, 10: 41,250ReAct, full replay, 20: 126,000ReAct + compaction, 2: 5,850ReAct + compaction, 5: 14,400ReAct + compaction, 10: 28,650ReAct + compaction, 20: 57,150ReAct + caching, 2: 2,745ReAct + caching, 5: 4,950ReAct + caching, 10: 9,525ReAct + caching, 20: 22,050plan-and-execute, 2: 5,210plan-and-execute, 5: 8,150plan-and-execute, 10: 13,050plan-and-execute, 20: 22,850ReAct, full replayReAct + compactionplan-and-executeReAct + cachingtool steps in the taskinput tokens per task (billed equivalent)
Billed-equivalent input tokens against tool steps under the planning numbers of ch3_cost.py: full replay, compaction, caching (at an assumed 10% price) and plan-and-execute.

At five steps the full-replay loop sends 15,750 input tokens and plan-and-execute 8,150; at twenty steps the ratio is 5.5 to one (126,000 against 22,850). With caching at our assumed 10% rate the twenty-step loop costs the equivalent of 22,050 tokens, about the same as the plan; with compaction it sends 57,150. So "ReAct is quadratic" is a statement about full replay at full price, not about ReAct. What a plan still buys under caching is fewer sequential model calls (latency) and an artefact a person can review before anything runs.

Loop (ReAct)Plan-and-execute
Good forfew steps, each depending on the last; exploratory tasksknown task shapes with several independent steps; plans a person should approve
Costsa model call per step; context growth under full replaya planner that must know the tools; replanning adds cost back
Fails byrepetition, drift, budget exhaustiona wrong plan executed in full before anyone notices
Measuresteps per task, repeat rate, tokens per taskdependency width (the parallel speed-up available), plan error rate, replanning rate
Research notes: Plan-and-Solve, ReWOO and LLMCompiler in detail (optional reading)

3.5 Checking the answer: who tells the agent it was wrong?

Row changed: verification.

Go back to the structured loop's miss: it looked up 8,849 m and calculated with 8,848. Could the agent itself have caught that? Two kinds of check are possible.

  • Ask the model to review its answer. It sees the same context and the same prior that produced 8,848. Nothing new enters the loop.
  • Check the answer against the evidence. A few lines of code can test that every number fed to the calculator appears in an earlier observation. 8,848 appears in no observation, so the check fails and the loop can say exactly what is wrong: "8,848 is not in any looked-up entry; Mount Everest's entry says 8,849."

The second check brings in information the model did not use. That difference is the whole of this section.

Reflection: try, check, critique, retry. The check decides whether it works.actorwrites an attemptevaluatorpass or failself-reflectionwhy it failed, in wordsmemoryreflections so farnext attempt, with the reflections in context (Reflexion, Figure 2)weak check: the model judges itselfstrong check: an external signal"review your answer and improve it"no new information enters the loopHuang et al.: accuracy fell on every benchmarkSelf-Refine: no gain on mathsuseful for style, format, checklist coverageunit tests, a compiler, a linter, a type checkera search result, a database row, a calculatora human review, a user's replyReflexion: HumanEval 80.1 to 91.0 with testsCRITIC: tools as critic; without them, no gainthe loop is the same; the signal is what you pay for. If there is no external signal,build one before you add a retry.
The reflection loop. Whether it helps depends on the check: the model judging itself, or an external signal such as tests, a compiler, a search or a person.

Two papers that disagree, and why they do not

Reflexion (Noah Shinn and colleagues, March 2023) let an agent retry after a failure with a short written reflection on what went wrong. On programming tasks the evaluator was a test suite the model wrote itself, run by a machine.

Reflexion's headline (GPT-4 on HumanEval, 80.1% to 91.0%) came with one honest loss: on MBPP the score fell from 80.1% to 77.1%, which the authors attribute to self-written tests that passed wrong code. Six months later Jie Huang and colleagues at Google DeepMind re-examined the self-correction results of 2023 and pointed out that several used an oracle: the true label decided when to stop retrying. Without it, asking a model to critique and revise its own reasoning made accuracy fall.

The two results fit together: Reflexion's gains came where a machine ran tests; Huang et al.'s losses came where the model had nothing but itself. Self-Refine (Madaan et al., 2023) found the same split from the other side: large gains on style and coverage tasks a model can see in its own text (dialogue, sentiment rewriting) and almost none on maths. CRITIC (Gou et al., 2023) supplied the constructive version: let the critic call tools (a search engine, a Python interpreter), and the gains come back; remove the tools and they mostly disappear. The research notes below give the tables.

The careful conclusion is conditional. Self-feedback can improve some outputs, especially where the defect is visible in the text, but it does not establish correctness. When correctness matters, use a check that is grounded in something independent of the model's own judgement.

Our experiment: repair with and without an external check

ch3_reflexion.py gives gpt-5-mini ten small coding tasks with precise specifications (an expression evaluator, a CSV line parser, a cache with expiry, a cron matcher, a slug generator, an edit distance, numbers to British English words, a topological sort, a strict Roman numeral parser, a semantic version parser). Each task is attempted five times, so there are 50 first attempts.

Each task has two test suites:

  • a feedback suite, the project's own tests, which a repair loop may run, show to the model and use to decide when to stop;
  • a held-out suite: different inputs for the same rules in the specification, never shown to the model and never used to stop anything. Every number below is measured on the held-out suite.

Both suites were first checked against reference solutions written for the script (all 20 pass), so a failure means the code is wrong, not the test.

Then every first attempt, right or wrong, goes through two repair strategies, so we can count fixes and breakages:

  • self-review: two rounds of "review your function against the specification; if you find a mistake, reply with the corrected code, otherwise reply with the same code unchanged". No test is run or shown.
  • feedback-test loop: run the feedback suite; if it fails, show the first failing test and ask for a fix; up to two more attempts; stop as soon as the feedback suite passes.

Terminal output of ch3_reflexion.py: the test suites checked against reference solutions, 20 of 20; a per-task table of held-out passes out of five for the first attempt, after self-review and after the feedback-test loop; and the summary: first attempts 43 of 50 pass the held-out tests; self-review 43 of 50 with one wrong-to-right and one right-to-wrong, the code changed in 45 of 50 attempts; the feedback-test loop 48 of 50 with five wrong-to-right and none right-to-wrong, six attempts entered the loop; the feedback suite accepted 44 attempts of which one failed held-out; 160 calls, 0.2262 dollars

50 first attempts at 10 coding tasks, graded on held-out testsfirst attemptfirst attempt: 86%86%after 2 rounds of self-reviewafter 2 rounds of self-review: 86%86%after the feedback-test loopafter the feedback-test loop: 96%96%all 50 first attempts went through both strategies, the right ones too. Self-review changed thecode in 45 of 50 attempts; it fixed 1 and broke 1. The feedback loop only touched the 6attempts its tests rejected; it fixed 5 and broke 0. Its check was not perfect: 1 of the 44 attemptsthe feedback suite accepted failed the held-out tests. 160 API calls, $0.23, gpt-5-mini.
Fifty first attempts graded on held-out tests: self-review left the pass rate at 86% (one fix, one breakage); the feedback-test loop raised it to 96% (five fixes, none broken).
Held-out pass rateWrong to rightRight to wrong
First attempt43/50 (86%)
After two rounds of self-review43/50 (86%)11
After the feedback-test loop48/50 (96%)50

What happened, concretely:

  • Self-review was busy but not useful. It changed the code in 45 of 50 attempts (38 already right), yet the pass rate did not move: it fixed one cron matcher and broke one correct cache.
  • The feedback loop only touched what its check rejected. Six attempts failed the feedback suite. Four of them were caches whose get restarted the expiry time, which the specification forbids (only set restarts it). The failing test read the entry at 9.9 seconds and expected it gone at 10; with that in hand the model fixed three of the four. The other fixes were an expression evaluator (unary minus with powers) and a Roman numeral parser that accepted "XCX".
  • The check was not perfect. The feedback suite accepted 44 first attempts, and one of those failed the held-out tests (a cron matcher that rejected the range "5-7" in the day-of-week field). The loop never looked at it again, so it stayed wrong. A check's false accepts cap any loop built on it.

The retry formula, and what it assumes

If each attempt passes with probability pp and a verifier can tell right from wrong, the chance that at least one of kk attempts passes is

P(success within k)=1−(1−p)kP(\text{success within } k) = 1 - (1 - p)^k

With p=0.5p = 0.5 and k=3k = 3 that is 88%. The expected number of attempts is 1−(1−p)kp\frac{1-(1-p)^k}{p}, and dividing expected cost by the success probability gives a cost per correct answer of c/pc/p for any kk.

This formula assumes three things, and real loops often break all three:

  1. Independent attempts. A feedback loop deliberately makes attempt 2 depend on attempt 1; that is the point of showing the failure.
  2. Constant success probability and cost. The second attempt's chance differs from the first's (better with good feedback, sometimes worse), and its context, so its cost, is larger.
  3. A perfect verifier. Real checks accept some wrong answers and reject some right ones, as our feedback suite did.

Treat the formula as a rough guide for independent resampling with a reliable check, and measure the real curve.

What an acceptance is worth

A verifier has two error rates. Let

  • pp = the probability that a candidate is correct,
  • rr = the probability the verifier accepts a correct candidate,
  • ff = the probability the verifier accepts a wrong candidate (the false-accept rate).

The probability that an accepted answer is correct follows from Bayes' rule:

P(correct∣accepted)=p rp r+(1−p) fP(\text{correct} \mid \text{accepted}) = \frac{p\,r}{p\,r + (1-p)\,f}

Worked example: p=0.9p = 0.9, r=1r = 1, f=0.2f = 0.2 gives 0.90.9+0.1×0.2=0.90.92≈97.8%\frac{0.9}{0.9 + 0.1 \times 0.2} = \frac{0.9}{0.92} \approx 97.8\%. A false-accept rate of 20% does not cap precision at 80%: precision depends on how many wrong candidates reach the verifier in the first place. The same verifier with p=0.3p = 0.3 gives 0.30.3+0.7×0.2≈68.2%\frac{0.3}{0.3 + 0.7 \times 0.2} \approx 68.2\%. An acceptance means less on hard tasks, where most candidates are wrong.

Terminal output of ch3_cost.py part five: the formula P of correct given accepted equals p r over p r plus one minus p times f, a table for p of 0.9, 0.7, 0.5 and 0.3 with three verifiers, including 97.8 percent for p 0.9, r 1 and f 0.2, and 68.2 percent for p 0.3; then best-of-n returning the first accepted candidate for p 0.3 and 0.7 and n of 1, 3, 5 and 10, showing that the chance of returning something rises with n while the precision of what is returned stays at 68.2 and 92.1 percent

02550751001030507090f = 0.05, r = 1, 10: 69f = 0.05, r = 1, 20: 83f = 0.05, r = 1, 30: 90f = 0.05, r = 1, 40: 93f = 0.05, r = 1, 50: 95f = 0.05, r = 1, 60: 97f = 0.05, r = 1, 70: 98f = 0.05, r = 1, 80: 99f = 0.05, r = 1, 90: 99f = 0.2, r = 1, 10: 36f = 0.2, r = 1, 20: 56f = 0.2, r = 1, 30: 68f = 0.2, r = 1, 40: 77f = 0.2, r = 1, 50: 83f = 0.2, r = 1, 60: 88f = 0.2, r = 1, 70: 92f = 0.2, r = 1, 80: 95f = 0.2, r = 1, 90: 98no verifier, 10: 10no verifier, 20: 20no verifier, 30: 30no verifier, 40: 40no verifier, 50: 50no verifier, 60: 60no verifier, 70: 70no verifier, 80: 80no verifier, 90: 90f = 0.05, r = 1f = 0.2, r = 1no verifiershare of candidates that are correct, p (%)P(correct | accepted), %p r / (p r + (1 - p) f)p = 90%, f = 0.2: 97.8%p = 30%, f = 0.2: 68.2%
P(correct | accepted) against the share of correct candidates, for two verifiers and none. The worked example, p = 90% and f = 0.2, gives 97.8%.
Research notes: Reflexion, Self-Refine, CRITIC and the self-correction debate (optional reading)

The paper's other results follow the same rule. On ALFWorld the evaluator was a heuristic that detected repeated actions and over-long trajectories, external to the model, and success rose by about 22 points over twelve trials. On HotpotQA the evaluator was exact match against the gold answer, which is an oracle; Huang et al. later argued that most of the gain on such tasks comes from the oracle rather than the reflection.

02550751001235p = 30% per attempt, 1: 30p = 30% per attempt, 2: 51p = 30% per attempt, 3: 66p = 30% per attempt, 5: 83p = 50% per attempt, 1: 50p = 50% per attempt, 2: 75p = 50% per attempt, 3: 88p = 50% per attempt, 5: 97p = 70% per attempt, 1: 70p = 70% per attempt, 2: 91p = 70% per attempt, 3: 97p = 70% per attempt, 5: 100p = 70% per attemptp = 50% per attemptp = 30% per attemptattempts allowed (k)P(success within k), %P = 1 - (1 - p)^k; assumesindependent attempts, fixed p,and a perfect verifier
Probability of success within k attempts for per-attempt pass rates of 30, 50 and 70 percent, under the formula's assumptions: independent attempts, the same pass rate on every attempt and a perfect verifier.

Row changed: verification and selection.

Without tools, the three direct answers to the Porto and Leeds question were "6 years", "9 years" and "5": a vote would only pick the most popular invention. On Everest, chain of thought gave 2,953 three times: a vote would be unanimous and wrong. Choosing among candidates helps only when the candidates contain the right answer and something can recognise it.

With the tools, the picture changes. Run the structured loop five times on the depot question and keep an answer only if a check passes (the evidence entries exist, every number used appears in an observation, the arithmetic recomputes). Now the selection has a real signal.

The idea predates the agent literature. Cobbe and colleagues at OpenAI trained a model to judge solutions to grade-school maths problems and used it to pick among many samples.

Best-of-n has a subtle limit, shown in the second table of the verifier screenshot above. If you return the first accepted candidate, more candidates raise the chance of returning something, but under independent candidates the precision of what you return stays at p rp r+(1−p) f\frac{p\,r}{p\,r + (1-p)\,f}, whatever n is. Two things make it worse in practice: correlated errors (the same wrong answer in many samples, which a lenient verifier then sees many times), and picking the top-scored candidate when the scorer has blind spots. Section 3.8 adds a further finding: how much extra sampling helps depends on how hard the question is.

Tree of Thoughts (Shunyu Yao and colleagues, May 2023) searched over partial solutions instead of complete ones: propose several next steps, score each ("sure", "maybe", "impossible"), expand only the promising ones.

Search over reasoning: propose, score, keep the best, repeat (b = 3, depth = 3)problemsuremaybeimpossibleanswerevery node scored sure, maybe or impossible by the modelwhat the tree costs (ch3_cost.py, breadth-first)per level: 3 propose calls + 9 evaluation callsthree levels: 36 model calls for one taskplanning numbers: 62,100 in + 4,590 out tokensabout 26x the cost of one chain of thoughtand 4.5x its latency (levels are sequential)the paper: Game of 24, GPT-4chain of thought 4%, best of 100 chains 49%,tree of thoughts b=5: 74% (Table 2);about 5.5k completion tokens per problem,close to 100 chain-of-thought trials (Table 7)worth it when steps can be scored before the end (a puzzle state, a test, a type check) andsingle chains fail early; not for open-ended writing or routine tool use
Tree of thoughts: propose several next steps, score them, expand the promising ones; with the call arithmetic of ch3_cost.py and the paper's Game of 24 numbers.

The task explains the gain: a small search space, cheap and fairly reliable scoring of partial states, and a first move that decides everything. The paper's own accounting puts the tree at about $0.74 per puzzle against $0.47 for a hundred chains (2023 GPT-4 prices), so per solved puzzle the two cost about the same ($1.00 against $0.96). On a task a single chain solves 90% of the time, the same tree would mostly be waste. LATS (Zhou et al., 2023) extended the search to actions and environment feedback; its results, read in the notes, came with an oracle setup and around 67 model calls per question.

One task, seven patterns: cost and latency relative to a direct answer (ch3_cost.py)cost per task (x direct)latency per task (x direct)direct answerdirect answer: 1.0x1.0xchain of thoughtchain of thought: 2.0x2.0xReAct, 5 tool stepsReAct, 5 tool steps: 12.3x12.3xplan-and-execute, 5 parallel stepsplan-and-execute, 5 parallel steps: 7.4x7.4xtree of thoughts, b=3, d=3tree of thoughts, b=3, d=3: 51.5x51.5xbest-of-5, model verifierbest-of-5, model verifier: 15.8x15.8xbest-of-5, free checkerbest-of-5, free checker: 9.8x9.8x: 1.0x1.0x: 3.5x3.5x: 12.5x12.5x: 6.2x6.2x: 15.5x15.5x: 4.4x4.4x: 3.5x3.5xthe tree is the outlier on both axes; best-of-n is cheap in time (parallel) and expensive in tokens;plan-and-execute buys back latency by running steps in parallel. None of these numbers says whichpattern is right: divide each by its accuracy first (Section 3.9).
Cost and latency of seven patterns relative to a direct answer, under the planning numbers of ch3_cost.py. Divide by accuracy before deciding.

What to measure before choosing among candidates: the single-attempt pass rate (near 0 or near 1, selection cannot help), the verifier's false-accept rate, and, for a tree, the step at which failed chains go wrong and the nodes expanded per solved task.

Research notes: Tree of Thoughts and LATS in detail (optional reading)

3.7 Changing the action: code, interfaces and fixed pipelines

Rows changed: action representation, and control across calls.

The depot question took the structured loop four look-up turns. If the action could be a short program, one turn would do:

python
trucks = {d: lookup_entry(d)['trucks'] for d in ['Rotterdam depot', 'Gdansk depot', 'Porto depot', 'Leeds depot']}
best = max(trucks, key=trucks.get)
print(best, trucks[best] - trucks['Leeds depot'])          # Rotterdam depot 79

The model writes this once; a sandbox runs it and returns one line. (This is an illustration; the chapter's scripts do not run model-written code.) That is the idea behind code as action.

JSON tool calls: three round tripsone code action in a sandboxlookup_rate("USD", "EUR")resultlookup_price("phone", "DE")resultconvert(price, rate)resulteach result passes through the modelbefore the next call can be maderate = lookup_rate("USD", "EUR")for country in ["DE", "FR", "IT"]: p = lookup_price("phone", country) prices[country] = convert(p, rate)print(min(prices.items(), key=...))one round trip; loops and conditions are free;the model reads one line back.CodeAct: up to 20 points higher success, fewer turnsthe counterpoint: Agentless, a fixed three-phase pipeline, no agent at alllocalisefiles > functions > linesrepairsample many patchesvalidatereproduce, test, rankSWE-bench Lite, mid-2024: 32.00% of issues resolved at $0.70 each,above every open-source agent of the time (Table 1)
Code as action against JSON tool calls, and the Agentless counterpoint: a fixed pipeline of localisation, repair and validation.

Two other results from 2024 belong next to CodeAct, because together they say what to change first.

The interface matters. SWE-agent (John Yang and colleagues at Princeton) gave a model a purpose-built set of commands for working in a code repository: a file viewer that shows a window of lines, a search that summarises its matches, an edit command that checks syntax. Changing one interface element at a time moved the resolved rate by several points with the model held fixed.

In its ablations a search command that summarises matches was worth six points of resolved rate over one that lists them all, and showing a 100-line window instead of whole files was worth about five. The same paper's behavioural analysis produced a phrase worth keeping: agents succeed quickly and fail slowly. Successful runs finished at a median of 12 steps and $1.21; unsuccessful ones averaged 21 steps and $2.52. A long run is more likely lost than about to succeed, which is an argument for setting the step budget from successful runs.

Sometimes no agent is needed. Agentless (Chunqiu Steven Xia and colleagues, July 2024) fixed GitHub issues with a fixed three-phase pipeline (find the location, sample many patches, validate with tests) in which the model never chooses the next step.

Read the three together. CodeAct changed the action format; SWE-agent changed what the tools return; Agentless removed model-chosen control and kept sampling plus verification. None says "agents are good" or "agents are bad". When the steps of a task are known, a fixed pipeline with a check can be cheaper and more reliable; when the model must choose, give it a good interface and, where composition matters, the option to write code. Six months after Agentless, Anthropic reported 49% on SWE-bench Verified with a minimal scaffold in which the model chose every step, with a stronger model. The right amount of fixed structure appears to move with model capability and with how verifiable the task is.

Research notes: CodeAct, SWE-agent, Agentless and the SWE-bench scaffolds in detail (optional reading)

3.8 What newer research changed (2024 to 2025)

Most papers above used 2022 and 2023 models. Three later results qualify their conclusions. None of them overturns the practical advice; each changes how strongly it should be stated.

Self-correction can be trained. Huang et al. showed that prompted models of 2023 got worse when asked to correct themselves. Aviral Kumar and colleagues at Google DeepMind (SCoRe, September 2024) trained a model with multi-turn reinforcement learning on its own correction attempts. Their table measures exactly the two quantities of our Section 3.5 experiment: wrong-to-right (Δi→c\Delta^{i \to c}) and right-to-wrong (Δc→i\Delta^{c \to i}).

The best use of extra compute depends on difficulty. Charlie Snell and colleagues (August 2024) compared ways of spending test-time compute: sampling many answers and searching with a verifier against letting the model revise its answer in sequence.

Reasoning models' chains are more faithful, but far from fully. Yanda Chen and colleagues at Anthropic (2025) repeated the Turpin-style test on reasoning models: insert a hint into the prompt, find cases where the hint changed the answer, and check whether the chain of thought mentions the hint.

3.9 Patterns in production

The papers measured benchmarks. Engineering write-ups describe systems with users. This section keeps two things apart: what the company reports (its own numbers, quoted, attributed to the exact system they describe) and our reading of which patterns the system uses. Where a page gives no number, none is given here. The full excerpts are in the research notes at the end of the section.

Patterns the companies describe in their own write-ups (filled = stated on the page)fixed pipelineReAct loopplan+executereflect on testssearch/best-of-njudge/human gateUber Genie, uReview, FixrLeakUber Genie, uReview, FixrLeak: fixed pipelineUber Genie, uReview, FixrLeak: ReAct loopUber Genie, uReview, FixrLeak: plan+executeUber Genie, uReview, FixrLeak: reflect on testsUber Genie, uReview, FixrLeak: search/best-of-nUber Genie, uReview, FixrLeak: judge/human gateAmazon Q code transformationAmazon Q code transformation: fixed pipelineAmazon Q code transformation: ReAct loopAmazon Q code transformation: plan+executeAmazon Q code transformation: reflect on testsAmazon Q code transformation: search/best-of-nAmazon Q code transformation: judge/human gateSpotify HonkSpotify Honk: fixed pipelineSpotify Honk: ReAct loopSpotify Honk: plan+executeSpotify Honk: reflect on testsSpotify Honk: search/best-of-nSpotify Honk: judge/human gateDoorDash support, AssistantDoorDash support, Assistant: fixed pipelineDoorDash support, Assistant: ReAct loopDoorDash support, Assistant: plan+executeDoorDash support, Assistant: reflect on testsDoorDash support, Assistant: search/best-of-nDoorDash support, Assistant: judge/human gateLinkedIn SQL Bot, Hiring Asst.LinkedIn SQL Bot, Hiring Asst.: fixed pipelineLinkedIn SQL Bot, Hiring Asst.: ReAct loopLinkedIn SQL Bot, Hiring Asst.: plan+executeLinkedIn SQL Bot, Hiring Asst.: reflect on testsLinkedIn SQL Bot, Hiring Asst.: search/best-of-nLinkedIn SQL Bot, Hiring Asst.: judge/human gateStripe MinionsStripe Minions: fixed pipelineStripe Minions: ReAct loopStripe Minions: plan+executeStripe Minions: reflect on testsStripe Minions: search/best-of-nStripe Minions: judge/human gateAirbnb test migrationAirbnb test migration: fixed pipelineAirbnb test migration: ReAct loopAirbnb test migration: plan+executeAirbnb test migration: reflect on testsAirbnb test migration: search/best-of-nAirbnb test migration: judge/human gateAnthropic SWE-bench scaffoldAnthropic SWE-bench scaffold: fixed pipelineAnthropic SWE-bench scaffold: ReAct loopAnthropic SWE-bench scaffold: plan+executeAnthropic SWE-bench scaffold: reflect on testsAnthropic SWE-bench scaffold: search/best-of-nAnthropic SWE-bench scaffold: judge/human gateGoogle Deep ResearchGoogle Deep Research: fixed pipelineGoogle Deep Research: ReAct loopGoogle Deep Research: plan+executeGoogle Deep Research: reflect on testsGoogle Deep Research: search/best-of-nGoogle Deep Research: judge/human gateour reading of the write-ups, not a survey: most rows have a verifier (tests, CI, a judgemodel) or a human gate. Klarna is left out: its assistant's architecture was not published.
The patterns companies describe in their own write-ups, as we read them. Klarna is left out because its architecture was not published.
Company, systemWhat the company reportsOur reading of the pattern
Uber, Genie (security and privacy channels)an LLM judge against a golden set cut evaluation from weeks to minutes; a relative 27% more acceptable answers and 60% less incorrect advicea fixed retrieval pipeline; progress came from cheap evaluation, not a new loop
Amazon, Q Developer Java upgradestens of thousands of applications moved to Java 17; developers review the plan before it runsplan-and-execute with a human approving the plan
Spotify, Fleet Management and HonkFleet Management automates about half of Spotify's pull requests since mid-2024; Honk, the coding agent inside it, has produced 1,500+ merged pull requests; a judge vetoes about a quarter of sessionsa home-made loop replaced by a coding agent with few tools, a verifiable goal and a judge
DoorDash, Dasher supporta two-tier guardrail cut hallucinations by 90% and severe compliance issues by 99%; a stronger guardrail was too slow and costlya fixed pipeline with an external critic and a human fallback
LinkedIn, SQL Botvalidators check tables exist and run EXPLAIN; errors go to a correction step with toolsreflection on an external signal (the database)
Stripe, Minionsat most two rounds of CI per run, because CI costs tokens, compute and timea state machine of fixed steps and agent steps; tests as the critic, capped
Airbnb, test migration75% of 3,500 files migrated in four hours, 97% after four days of tuning, the rest by handa fixed per-file pipeline with validation errors fed back and capped retries

Spotify: one system inside another

Spotify's numbers are easy to misattribute. "Around half of Spotify's pull requests" describes Fleet Management, the system that applies automated changes across thousands of repositories and existed before any model was involved. Honk, the coding agent added inside it, accounts for the title's "1,500+" merged AI-generated pull requests. The team reports that its first home-made loop "tended to get lost when it filled up its context window", and that its LLM judge, which vetoes about a quarter of sessions, has not yet been evaluated. Our reading: context growth and a missing check broke the first loop, and a judge is a check of unknown precision until it is measured.

Klarna: cost savings did not settle the quality question

In February 2024 Klarna reported that its AI assistant handled two-thirds of customer service chats in its first month, "the equivalent work of 700 full-time agents". In May 2025, as Fortune reported from a Bloomberg interview, its chief executive said that "cost unfortunately seems to have been a too predominant evaluation factor" and that the result was "lower quality", and Klarna began hiring people for customer service again. The architecture was never published, so nothing here says which pattern it used. What the story shows is about measurement: reporting cost and volume first left the quality question open, and it had to be answered later.

Research notes: the company write-ups in detail (optional reading)

Uber. Three systems cover most of this chapter's patterns.

Amazon.

Spotify.

DoorDash.

DoorDash's later posts trace the same team moving from fixed workflows to single agents and then to several agents. A November 2025 post names the single agent's limit as "context pollution", and a June 2026 post on the consumer DoorDash Assistant reports that "the largest potential production-failure category is grounding" and that "the fix in each case has been to route the agent's claim through a tool call against the system of record".

LinkedIn.

Stripe.

Airbnb. The March 2025 post "Accelerating Large-Scale Test Migration with LLMs" (the Airbnb Tech Blog blocks automated screenshots, so it is quoted) describes migrating "nearly 3.5K React component test files" between testing frameworks, a job estimated at "1.5 years of engineering time" and finished "in just 6 weeks". Each file went through a fixed pipeline ("we modeled this flow like a state machine") that would "retry steps multiple times until they passed or we reached a limit", with "the validation errors and the most recent version of the file" in each retry prompt. The team reports "75% of our target files in just four hours", "from 75% to 97%" after four days of tuning, and the "remaining 3%" finished by hand. Our reading: a fixed pipeline, reflection on an external signal, capped retries and a human tail.

Anthropic and Google.

3.10 Choosing a pattern

Back to Northwind. Suppose the company wants an assistant that answers staff questions from its encyclopaedia. Here is how this chapter's measurements would guide the build, one decision at a time.

  1. A labelled set and the simplest system. Twelve typed questions, an exact grader. Direct answering: 3%, failing on missing facts.
  2. Missing facts call for tools in a loop. Plain ReAct: 56%; the traces show repetition and tool misses.
  3. Fix the loop and the tool first. Guard plus improved tool: 100%. A structured loop: 94%, with misses an evidence check would catch.
  4. Add the check the misses call for: an evidence check costs a few lines and no model calls.
  5. A plan only if steps are many and independent, priced with caching and compaction included.
  6. Best-of-n or search only if single attempts are middling and a reliable check exists.

That sequence is a starting heuristic, not a law. The five dimensions of Section 3.1 are independent, so you can skip or combine steps when the task makes the need obvious: code with a test suite can start at "reflection on tests"; a puzzle with cheap state scoring can start at search; a task whose steps are fully known can start, and stay, as a fixed pipeline.

A starting heuristic: begin simple, add one thing per measured failurestartyesone call, or a fixed pipeline if the steps are knownmeasure accuracy, cost and latency on a labelled setno, or still failingwrong on multi-step reasoning?yesadd chain of thought, or a reasoning modelcost: output tokens; do not assume faithfulnessno, or still failingfacts or state not in the model?yesadd tools in a ReAct loop: step budget, repeat guardcost: n calls on a growing contextno, or still failingmany steps; latency or tokens too high?yesplan first, execute with small parallel contextscost: a wrong plan is found late; replanno, or still failingan external check exists?yesadd reflection on tests, compiler, search, humancost: k attempts; self-review alone is weakno, or still failingearly choices decide; states scorable?yessearch: best-of-n with a verifier, or a treecost: 10x to 100x tokens; puzzles, code with testsnot a fixed sequence: skip or combine rungs when the task makes the need obvious (code with testsstarts at reflection); at every arrow re-measure accuracy, cost per correct answer and p95 latency
A starting heuristic, not a fixed sequence: add one thing per measured failure, skip or combine steps when the task makes the need obvious, and re-measure.
Task propertyHow to measure itWhat it argues for
Verifiability: can an answer be checked without knowing it?list the available checks and measure their false-accept rateswith a strong check: reflection on it, and best-of-n; without one: short loops and a person at the end
Step countthe distribution of tool calls in solved runsone to three: a direct call or a short loop; more, with independent steps: a plan
Knowledgethe direct-answer baselinehigh: chain of thought may be enough; low: tools, and check that tools do not hurt the easy cases
Latency budgetthe product's p95 against the pattern's sequential depthtight: parallel candidates or a plan with parallel steps
Cost of a wrong answerwhat happens downstreamhigh: spend on verification; low: the cheapest pattern that meets the bar
Shape of failuresthe step at which failed runs went wrongrepetition: a guard; tool misses: better observations; early fatal choices: search; late, checkable errors: reflection

ch3_cost.py puts the cost per correct answer of seven patterns side by side under planning numbers:

Terminal output of ch3_cost.py parts one and two: a table of seven patterns with calls, input and output tokens, latency and cost at two prices; relative cost and latency against a direct answer; and the cost per correct answer, with measured accuracies for direct answer 3 percent, chain of thought 22 percent and the plain ReAct loop 56 percent, and assumed accuracies for the others

Read the second table with care. The three accuracies marked "mea" are measured (Section 3.3, the plain loop); the rest are assumptions to show the arithmetic. Under those numbers, at frontier-model prices, chain of thought costs about 4.4 cents per correct answer and the plain loop about 10.9 cents, because almost half its runs were wasted; a plan assumed to reach 80% would cost about 4.6 cents and a tree assumed to reach 90% about 28 cents. Put in your own accuracies and the ranking can change; the formula will not.

The one rule of the chapter, stated as a condition: add a pattern when a measured failure calls for it, and keep it only if accuracy, cost per correct answer and latency still justify it after you re-measure.

Exercises

  1. Your wiki agent uses a ReAct loop with a ten-step budget, and 30% of failed runs hit the budget. Design three measurements that separate repetition, unhelpful tool misses and genuinely long tasks, and name the fix each one points to.
  2. A team wants to add "the model reviews its answer before replying" to a customer-facing agent. Using Huang et al., SCoRe's Table 2 and our repair experiment, write the one-paragraph case for measuring wrong-to-right and right-to-wrong changes first, and list three external checks available in a typical support system.
  3. Re-run ch3_cost.py with your own prompt size, observation size, prices and cache discount. At what step count does plan-and-execute become cheaper than the loop with full replay, and does it still win with caching?
  4. Your verifier accepts 5% of wrong answers and all right ones. Compute P(correct∣accepted)P(\text{correct} \mid \text{accepted}) for p=0.8p = 0.8 and p=0.2p = 0.2. Which of your task types sits nearer each value, and what does that imply about where to spend on a better verifier?
  5. Take the Spotify and Airbnb cases. For each, write one sentence on what the company reports and one on your reading of the pattern, and name one number you would want that neither post gives.

Key takeaways

  • A reasoning pattern is a choice on one or more of five independent dimensions: reasoning inside one call, control across calls, action format, verification and selection, and context management. They combine; they are not a ladder.
  • Chain of thought helps a model use what it knows (public questions: 6% to 44% in our run) and cannot add facts it lacks (private: 0% to 0%). Self-consistency returns the plurality answer; it fails when errors are correlated.
  • ReAct grounds the model in tool results but fails by repetition and unhelpful tool misses. In our ablation the plain loop scored 56%, the guard alone 75%, the improved tool alone 75%, and both 100%. Its "0% hallucination" is a share of sampled failed runs in one study, not an overall rate.
  • A modern structured loop (schema-checked tool calls, explicit state, budgets, typed final answer, trace) scored 34 of 36 on the same questions; its misses were a prior overriding an observation and a strict grader.
  • Plan-and-execute cuts sequential calls and, under full replay, tokens; caching and compaction change the token comparison, so state the assumptions.
  • Reflection is worth what its check is worth. In our held-out experiment, self-review left 43 of 50 unchanged in total (one fix, one breakage); the feedback-test loop reached 48 of 50 (five fixes, none broken).
  • A verifier's false-accept rate is not its precision: p=0.9p = 0.9, r=1r = 1, f=0.2f = 0.2 gives about 97.8% precision, but p=0.3p = 0.3 gives 68.2%. More candidates raise the chance of returning something, not its precision.
  • Newer work qualifies the 2023 results: self-correction can be trained (SCoRe), the best use of test-time compute depends on difficulty (Snell et al.), and reasoning models' chains are more faithful but often still omit what drove the answer (Chen et al.).
  • In production write-ups, keep a company's reported numbers apart from your reading of its pattern.

References

Papers

  1. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, Denny Zhou. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022 (arXiv January 2022).
  2. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR 2023 (arXiv March 2022).
  3. Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, Yusuke Iwasawa. Large Language Models are Zero-Shot Reasoners. NeurIPS 2022 (arXiv May 2022).
  4. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, John Schulman. Training Verifiers to Solve Math Word Problems. arXiv October 2021.
  5. OpenAI. OpenAI o1 System Card. arXiv December 2024.
  6. DeepSeek-AI. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv January 2025 (version 1; a revised version appeared in Nature, 2025).
  7. Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. NeurIPS 2023 (arXiv May 2023).
  8. Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, Denny Zhou. Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024 (arXiv October 2023).
  9. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023 (arXiv October 2022).
  10. Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, Ee-Peng Lim. Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models. ACL 2023 (arXiv May 2023).
  11. Binfeng Xu, Zhiyuan Peng, Bowen Lei, Subhabrata Mukherjee, Yuchen Liu, Dongkuan Xu. ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models. arXiv May 2023.
  12. Sehoon Kim, Suhong Moon, Ryan Tabrizi, Nicholas Lee, Michael W. Mahoney, Kurt Keutzer, Amir Gholami. An LLM Compiler for Parallel Function Calling. ICML 2024 (arXiv December 2023).
  13. Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao. Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023 (arXiv March 2023).
  14. Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, Peter Clark. Self-Refine: Iterative Refinement with Self-Feedback. NeurIPS 2023 (arXiv March 2023).
  15. Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, Weizhu Chen. CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. ICLR 2024 (arXiv May 2023).
  16. Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik Narasimhan. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. NeurIPS 2023 (arXiv May 2023).
  17. Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, Yu-Xiong Wang. Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models. ICML 2024 (arXiv October 2023).
  18. Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, Heng Ji. Executable Code Actions Elicit Better LLM Agents. ICML 2024 (arXiv February 2024).
  19. John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, Ofir Press. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024 (arXiv May 2024).
  20. Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, Lingming Zhang. Agentless: Demystifying LLM-based Software Engineering Agents. arXiv July 2024.
  21. Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal Behbahani, Aleksandra Faust. Training Language Models to Self-Correct via Reinforcement Learning. arXiv September 2024.
  22. Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv August 2024.
  23. Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Samuel R. Bowman, Jan Leike, Jared Kaplan, Ethan Perez. Reasoning Models Don't Always Say What They Think. arXiv May 2025; research post Reasoning models don't always say what they think, Anthropic, 3 April 2025.

Engineering blogs and docs

  1. Anthropic. Building effective agents. Research blog, 19 December 2024.
  2. Anthropic. Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet. Engineering blog, 6 January 2025.
  3. Anthropic. Code execution with MCP: Building more efficient agents. Engineering blog, 4 November 2025.
  4. Google (Dave Citron). Try Deep Research and our new experimental model in Gemini, your AI assistant. The Keyword, 11 December 2024.
  5. Uber. Enhanced Agentic-RAG: What If Chatbots Could Deliver Near-Human Precision?. Engineering blog, May 2025.
  6. Uber. uReview: Scalable, Trustworthy GenAI for Code Review at Uber. Engineering blog, 12 August 2025.
  7. Uber. FixrLeak: Fixing Java Resource Leaks with GenAI. Engineering blog, May 2025.
  8. AWS (Aytul Arisoy Cholkar). Amazon Q Developer just reached a $260 million dollar milestone. AWS DevOps blog, 1 August 2024.
  9. Spotify (Max Charas, Marc Bruggmann). 1,500+ PRs Later: Spotify's Journey with Our Background Coding Agent (Honk, Part 1). Engineering blog, 6 November 2025.
  10. Spotify. Background Coding Agents: Context Engineering (Honk, Part 2). Engineering blog, 24 November 2025.
  11. Spotify. Background Coding Agents: Predictable Results Through Strong Feedback Loops (Honk, Part 3). Engineering blog, 9 December 2025.
  12. DoorDash. Path to high-quality LLM-based Dasher support automation. Engineering blog, 17 September 2024.
  13. DoorDash. Beyond Single Agents: How DoorDash is building a collaborative AI ecosystem. Engineering blog, 11 November 2025.
  14. DoorDash. Building DoorDash Assistant: An engineering overview. Engineering blog, 11 June 2026.
  15. LinkedIn (Albert Chen and colleagues). Practical text-to-SQL for data analytics. Engineering blog, 9 December 2024.
  16. LinkedIn (Xiaoyang Gu, Xie Lu, Daniel Hewlett). Building the agentic future of recruiting: how we engineered LinkedIn's Hiring Assistant. Engineering blog, 21 October 2025.
  17. Stripe (Alistair Gray). Minions: Stripe's one-shot, end-to-end coding agents and Part 2. stripe.dev blog, 9 and 19 February 2026.
  18. Airbnb (Charles Covey-Brandt). Accelerating Large-Scale Test Migration with LLMs. Airbnb Tech Blog, March 2025.
  19. Klarna. Klarna AI assistant handles two-thirds of customer service chats in its first month. Press release, 27 February 2024; and Irina Ivanova, Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver, Fortune, 9 May 2025, reporting the Bloomberg interview of 8 May 2025.

Code for this chapter

  1. code/agents/ch3_react.py: twelve two-hop questions, six system configurations (including the four-way ablation of the guard and the improved tool), three repeats, a typed exact grader with a self-test; results in results/ch3_react.json and the two ch3_react*_stdout.txt logs.
  2. code/agents/ch3_structured_loop.py: a complete structured-tool loop (native tool calls, schema validation, explicit state, step and token budgets, error handling, a JSON-lines trace); results in results/ch3_structured_loop.json and results/ch3_structured_trace.jsonl.
  3. code/agents/ch3_reflexion.py: ten coding tasks with feedback and held-out test suites (both checked against reference solutions), self-review against a feedback-test loop on every first attempt; results in results/ch3_reflexion.json.
  4. code/agents/ch3_cost.py: the cost and latency arithmetic, cost per correct answer, the retry formula and its assumptions, token growth with caching and compaction, and the verifier precision and best-of-n tables; results in results/ch3_cost.json.
  5. code/agents/figs_ch3.py and code/agents/shots_ch3.py: the figures (with an automatic overlap and clipping check) and the paper excerpts.

Next

→ Chapter 4: Workflow patterns

This chapter kept one agent and changed how it thinks, acts and checks. Chapter 4 keeps the thinking and changes the plumbing around it: workflow patterns in which code, not the model, decides the path. Prompt chaining, routing, parallel fan-out, orchestrator and workers, and evaluator and optimiser are ways of arranging several model calls so that each is small, checkable and cheap. Several systems in this chapter (Uber's review pipeline, Stripe's state machine, Airbnb's migration) are workflows with an agent inside one step, and Snell et al.'s finding that compute should follow difficulty is, in practice, a routing decision.