Agentic Systems · Part 2 · Patterns

Chapter 4 · Workflow patterns: chaining, routing, parallel work, orchestrators and evaluators

Workflow patterns: prompt chains, routers, parallel calls, orchestrators and evaluators; what each costs, how each fails and how to test it.

Goal: by the end of this chapter you can draw any LLM system as a graph of steps, say which steps your code decides and which the model decides, and choose among the five workflow patterns (prompt chaining, routing, parallelisation, orchestrator-workers, evaluator-optimiser) by what each costs, how it fails and how you will test it. One running example carries you through all five, four small experiments against a real model show where each pattern earns its keep, and one complete program puts them together with the plumbing production needs: typed state, schema checks, budgets, retries, checkpoints, a human approval step and a trace.


4.1 Workflows versus agents: who holds the control flow

The running example: Northwind Home's support inbox

Northwind Home is a fictional online shop that sells furniture and kitchenware in the UK. Its support inbox receives a few thousand messages a week: "order #48213 arrived with the screen cracked", "I was charged twice for ORD-10988", "do you ship to the Isle of Man?", "the sofa from order 61234 has a torn seam, I want my money back, £649". The team wants a model to read each message, work out what it is about, draft a reply, check the reply against company policy, and send it, with a person approving anything that costs money.

The first version is one prompt: "here is the policy, here is the message, write the reply". It mostly works, and it fails in ways nobody can see from outside: the order number is quoted wrongly, the policy's ban on promising delivery dates is ignored, a refund is offered on a question about shipping zones. Every fix is another sentence in a prompt that keeps growing, and every new sentence can break a case that used to work.

This chapter rebuilds that system five times, once per pattern, and asks the same four questions each time: what changes, what it costs, how it fails, and how you test it. Underneath all five is one question: who decides what happens next?

Anthropic's December 2024 post "Building effective agents" drew this line, and the industry adopted it. It names five workflow patterns, and Sections 4.2 to 4.6 take them in turn.

Who holds the control flow: from a fixed chain to a free loopchainfixed stepsin coderouterLLM picksone branchparallelfixedfan-outorchestratorLLM splitsthe workevaluatorLLM judgeswhen to stopagentLLM picksevery steppredictablepredictable, chain: 5 of 5predictable, router: 4 of 5predictable, parallel: 4 of 5predictable, orchestrator: 3 of 5predictable, evaluator: 3 of 5predictable, agent: 1 of 5easy to testeasy to test, chain: 5 of 5easy to test, router: 4 of 5easy to test, parallel: 4 of 5easy to test, orchestrator: 2 of 5easy to test, evaluator: 3 of 5easy to test, agent: 1 of 5cost, latencycost, latency, chain: 1 of 5cost, latency, router: 1 of 5cost, latency, parallel: 2 of 5cost, latency, orchestrator: 3 of 5cost, latency, evaluator: 3 of 5cost, latency, agent: 5 of 5novel inputsnovel inputs, chain: 1 of 5novel inputs, router: 2 of 5novel inputs, parallel: 2 of 5novel inputs, orchestrator: 4 of 5novel inputs, evaluator: 3 of 5novel inputs, agent: 5 of 5workflows: the engineer writes the pathagents: the model writes the pathBars are qualitative (0 to 5), a summary of this chapter's measurements and the papers it reads, nota benchmark.
Six arrangements of model calls, ordered by who decides the next step. On the left the engineer writes every path, and the system is predictable, easy to test, cheap and fast, but brittle on inputs nobody foresaw; on the right the model decides every step and the trade reverses. The five workflow patterns sit in between, each handing the model one more kind of decision. The bars are a qualitative summary, not a measurement.

Read the figure as a sequence of hand-overs. A chain hands the model no decisions about flow. A router hands it one, made once: which branch should this input take? Parallel sections hand it none; code fans the work out and merges it. An orchestrator hands it the shape of the work: how many subtasks, and which. An evaluator loop hands it the decision to stop. An agent hands it every step. Each hand-over buys adaptability and usually costs predictability, testability, money and time. So a useful default is to hand the model a decision only when you cannot write that decision in code, and to know which decisions you have handed over.

Graphs: the vocabulary of workflows

Every workflow can be drawn as a graph, and the frameworks that run them (LangGraph, the OpenAI Agents SDK, Temporal, AWS Step Functions) all use graph words.

A workflow as a graph: nodes, edges, a gate, a conditional edge, checkpointsSTARTextractgateschema ok?yesclassifyno: retry once with the errorrefundotherrefund flowreply flowhuman approvesEND= checkpoint (state saved)nodea step: an LLM call, a tool call or plain code; reads state, returns an updateedgea fixed transition: after this node, always run that oneconditional edgea function of the state picks the next node (here: refund or other)gatea programmatic check between steps; failing it retries or stops the chaincheckpointthe state written to storage after a node, so a run can pause and resume
Northwind's ticket flow drawn as a graph. An extract node feeds a gate that checks the output's schema and retries on failure; a classify node feeds a conditional edge that sends refunds to a flow with human approval and everything else to a reply flow. Squares mark checkpoints, where the state is saved so the run can pause or resume.
Research notes: OpenAI's view from the agent end (optional reading)

OpenAI's practical guide of April 2025 frames the same choice from the other side: start with one agent and add structure only when it fails.

The two vendors start at opposite ends and meet in the middle: add a moving part when a measurement asks for it.

Discussion

  1. Is a router an agent? The model makes a decision about flow. My view: no, as long as the branches are fixed in code and the router can be tested against labels like any classifier; it becomes agent-like when the model can name a branch you did not write.
  2. Does the distinction matter to users? They see an answer, not a graph. But latency ceilings, cost ceilings and testability are properties of the graph, and those are what break in production.
  3. Will better models make workflows obsolete? Better models move the line towards agents for open-ended tasks. For high-volume, well-understood tasks, a fixed path's predictability and cost per call are part of the product, so the line moves more slowly there.

4.2 Prompt chaining: fixed steps, with gates in between

The example: split Northwind's one prompt into three

The one-prompt system drops requirements because it is asked to do four things at once: understand the ticket, choose a response, write it, and obey the policy. A prompt chain gives each job its own call:

  1. Extract: read the ticket and return JSON with the issue type (one of five), the order id in the form ORD-NNNNN, and the urgency.
  2. Draft: from those facts, choose the reply category from a fixed table and write the reply.
  3. Check: test the reply against the policy (at most 70 words, quote the order id, no promises of timing such as "as soon as possible", no talk of refunds outside billing).

Between the steps sit gates: code that checks the previous step's output before the next one runs.

Prompt chaining: fixed steps, a programmatic gate after each one1extractticket -> JSON:type, order id,urgencygate2draftfacts -> categoryand a reply(JSON)gate3checkreply -> policyverdict(JSON)sendreplyfail: retry the step once, with the error in the promptfails twice: stop, hand to a personwhat a gate can check, for free and in milliseconds- the output parses as JSON- every required key is present- enum values in the allowed list- ids match a pattern (ORD-NNNNN)- word limits, banned phrases- a value exists in the databasewhat a gate cannot check: a plausible value that is wrong (a valid "delivery" label on a questionabout shipping zones). Those need a labelled set and an eval, not a schema.
Northwind's three-step chain: extract facts as JSON, draft a reply, check it against the policy. A gate after each step checks the output in code; a failed gate retries the step once with the error in the prompt, and a second failure hands the ticket to a person. Gates catch broken formats and impossible values; they cannot catch a plausible value that is wrong.

Why chains need gates: the arithmetic

A chain multiplies probabilities. If each step returns a usable, correct output with probability rr, the whole chain of nn steps succeeds with probability

P(chain)=r n.P(\text{chain}) = r^{\,n}.

Assumptions: steps fail independently, a bad output anywhere ruins the result, and no later step repairs an earlier mistake. Real chains break all three in both directions: one confusing ticket makes several steps fail together, and a good drafting step can paper over a slightly wrong extraction. Treat the formula as a way to see the shape, not a forecast. The shape is unforgiving: five steps at 90% succeed 59% of the time; ten succeed 35%.

Now add a gate that detects a fraction dd of bad outputs and retries the step once. The step succeeds if the first attempt is good, or if it is bad, the gate notices and the retry is good:

r′=r+(1−r) d r.r' = r + (1 - r)\, d\, r .

Assumptions: the retry succeeds with the same probability rr as the first attempt (in practice the error message often helps, so this is a floor), and the gate never rejects a good output. With r=0.9r = 0.9 and d=0.8d = 0.8, r′=0.972r' = 0.972, and five steps succeed 87% of the time instead of 59%, for about 8% more calls. The large condition is dd: the share of bad outputs the gate can see. A gate sees a missing key, an unknown label, a malformed id, a reply over the word limit. It cannot see a well-formed answer that is wrong.

End-to-end success of an n-step chain: r to the power n02550751001234567891099% per step, n=1: 9999% per step, n=2: 9899% per step, n=3: 9799% per step, n=4: 9699% per step, n=5: 9599% per step, n=6: 9499% per step, n=7: 9399% per step, n=8: 9299% per step, n=9: 9199% per step, n=10: 9090% per step, n=1: 9090% per step, n=2: 8190% per step, n=3: 7390% per step, n=4: 6690% per step, n=5: 5990% per step, n=6: 5390% per step, n=7: 4890% per step, n=8: 4390% per step, n=9: 3990% per step, n=10: 3590% + gate (d=0.8), n=1: 9790% + gate (d=0.8), n=2: 9490% + gate (d=0.8), n=3: 9290% + gate (d=0.8), n=4: 8990% + gate (d=0.8), n=5: 8790% + gate (d=0.8), n=6: 8490% + gate (d=0.8), n=7: 8290% + gate (d=0.8), n=8: 8090% + gate (d=0.8), n=9: 7790% + gate (d=0.8), n=10: 7599% per step90% + gate (d=0.8)90% per stepsteps in the chain (n)P(chain succeeds), %At 95% per step, ten steps succeed 60% of the time. A gate that detects 80% of bad outputs andretries once turns a 90%-per-step step into a 97% one: five such steps go from 59% to 87% end toend, for about 8% more calls (ch4_cost.py part 1).
End-to-end success of an n-step chain, from ch4_cost.py. At 99% per step a ten-step chain succeeds 90% of the time; at 90% per step, 35%. A gate that detects 80% of bad outputs and retries once lifts a 90% step to 97%, so a five-step chain goes from 59% to 87%. All three curves assume independent steps.

Terminal output of ch4_cost.py, the first parts: the chain-reliability table for per-step reliability 0.90, 0.95 and 0.99 with and without a gate that detects 80 percent of bad outputs, five steps at 90 percent giving 59 percent and 87 percent with the gate, with the assumptions printed; then routing cost against the share of hard inputs and the router's recall and false-escalation rate; then sectioning latency and the plurality-voting table with correlated errors

Our experiment: Northwind's chain, with and without gates

The script ch4_chain.py runs this chain on twenty short tickets with a real model (gpt-5-mini, reasoning effort set to minimal). The tickets are deliberately messy: order numbers written five ways ("order #48213", "ORD 55501", "order no. 30377"), one in Spanish, one mentioning two orders, several with no order at all. Each ticket has gold labels: its issue type and its canonical order id.

There are four conditions, varying one factor at a time. The step-1 prompt is either strict (it names the keys, the five allowed values and the ORD-NNNNN format) or loose ("extract the issue type, the order id and the urgency ... reply in JSON", the kind of prompt a first version often has). Each runs without gates or with gates (a schema check after step 1, a schema and policy check after step 2, one retry each). Success has two parts, reported separately because they differ in independence:

  • Gold correct: the right reply category and the gold order id quoted in the reply. No gate ever sees the gold labels, so this is an independent measure.
  • Policy correct: word limit, no timing promises, no refund talk outside billing. Gate 2 checks the same rules, so in the gated conditions this part is graded by the same code that gave the feedback. It shows compliance, not independent quality.

The whole run cost about four US cents; token counts come from the API's usage fields.

python
# code/agents/ch4_chain.py (abridged): one chain step with a gate. Without gates, whatever came back flows on.
def step(prompt, gate, gated, use):
    text, u = llm.ask(prompt, max_tokens=500)                 # one model call; usage recorded from the API
    use.append(u)
    d = llm.parse_json(text)
    if not gated:
        return d, text, 0
    problems = gate(d)                                         # e.g. "order_id must match ORD-NNNNN or be null, got '48213'"
    if not problems:
        return d, text, 0
    retry = prompt + '\n\nYour previous answer failed these checks:\n- ' + '\n- '.join(problems) + \
            '\nReply again with a corrected JSON object only.'
    text, u = llm.ask(retry, max_tokens=500)                   # one retry, with the exact failures in the prompt
    use.append(u)
    d = llm.parse_json(text)
    return (d if not gate(d) else None), text, 1               # None stops the chain: hand the ticket to a person

Terminal output of ch4_chain.py: the loose prompt without gates, where tickets fail at step 1 because the model returned free-text issue types such as Damaged item, cracked screen, and bare order numbers such as 48213; then the summary table: strict prompts 19 of 20 with and without gates, loose without gates 12 of 20, loose with gates 16 of 20 at 78 calls instead of 60, policy correct 19 or 20 of 20 in every condition, and the step-3 model checker flagging policy P1 71 times

Our chain on 20 tickets: where the gates earn their placeend-to-end successcalls per ticketstrict prompt, no gatesstrict prompt, no gates: 19/2019/20strict prompt, gatesstrict prompt, gates: 19/2019/20loose prompt, no gatesloose prompt, no gates: 12/2012/20loose prompt, gatesloose prompt, gates: 16/2016/20: 3.03.0: 3.03.0: 3.03.0: 3.93.9Strict prompt: the gates had nothing to catch, and the one failure (a shipping question labelled"delivery") was a valid label, so no gate could see it. Loose prompt: 8 of 20 chains broke onfree-form labels and bare order numbers; with gates 16 of 20 succeeded, for 18 extra calls. Tworetries "fixed" the order id by setting it to null, which the schema allows. The step-3 modelchecker flagged "over 70 words" 71 times across the four runs on replies a word count passed.
Results of the final run of ch4_chain.py. With a strict first-step prompt the gates had nothing to catch: 19 of 20 either way. With a loose prompt, 8 of 20 chains broke on free-form labels and bare order numbers; gates with one retry brought success to 16 of 20 for 18 extra calls. The model checker in step 3 disagreed with a plain word count on most tickets.

When the prompt is specific, the gates are idle. With the strict prompt the model returned valid labels and canonical ids on every ticket; the gates never fired, and both conditions scored 19 of 20. The one failure was "Do you ship to the Isle of Man?", labelled "delivery" instead of "other". That is a valid label, so no gate could catch it; it needs a labelled set to find.

When the prompt is loose, the gates do the work. With the loose prompt every step-1 output parsed as JSON and almost none was usable: issue types came back as "Damaged item / cracked screen" and "Incorrect VAT rate on invoice", order ids as "48213" and "ORD 55501". Without gates these flowed on and 8 of 20 replies went to the wrong category or failed to quote the order id. With gates, all twenty tickets failed their first check, were retried once with the exact problems listed, and 16 of 20 succeeded. Of the four failures, one was the Isle of Man label, one was a ticket the gate correctly stopped and handed to a person (it mentioned two orders), and two are the instructive ones: the retry "fixed" the order id by setting it to null, which the schema allows for tickets with no order. A retry optimises for the gate, so a gate with an easy wrong exit gets that exit.

A model is a poor checker of mechanical rules. Step 3, the model asked whether the reply broke the policy, disagreed with a plain-code check on most tickets. Across the four conditions it flagged "over 70 words" 71 times on replies that a word count passed. A rule you can write in code (a word count, a list of banned phrases, a regular expression) is better checked in code, and a model checker belongs where code cannot reach, such as tone, after you have measured it (Section 4.6).

The research behind chaining

The idea was first studied as a question of human control. In October 2021 Tongshuang Wu, Michael Terry and Carrie Cai at Google posted "AI Chains" (CHI 2022). Large models of the time did well on single operations and poorly on tasks with several parts, and users could not see or fix what went wrong inside one big prompt.

The study's measures were participants' ratings, not accuracy; the authors note the limits (twenty people, two tasks, a 2021 model). In 2022 two papers turned decomposition into a reasoning method: Least-to-Most prompting (Zhou et al., Google) and Decomposed Prompting (Khot et al., Allen Institute for AI). Both found the largest gains on problems with many steps, and both point at the same practical conclusion: the decomposition is the hard part, and when an engineer can write it, it can live in code.

Research notes: Least-to-Most, Decomposed Prompting and the AI Chains abstract (optional reading)

What a production team learned about gates

Prompt chaining in practice

One big promptPrompt chain with gates
Prosone call, lowest latency; nothing to wireeach step easy for the model and testable alone; different models per step; format failures caught in code; one step changes without touching the others
Consrequirements dropped silently; nothing to test between input and outputlatency adds up; errors propagate; interfaces between steps must be designed and versioned; a gate can push a retry into a wrong-but-valid answer
When to pick itthe task is one operation (classify, extract, rewrite)the task has distinct operations in a known order, especially if some can be code
Production failurea missed requirement nobody notices until a user doesan upstream change (new input formats) degrades every step downstream

What to measure before you decide

MeasurementHowWhat it tells you
Per-step success ratelabel each step's output on a hundred inputsthe chain's ceiling; which step to improve first
Gate detection rate ddlabel the failures the gate missedhow much a gate can recover; valid-but-wrong is the rest
First fault positionfor each failed run, the first step whose output was wrongwhether to fix extraction or drafting
Retry outcomefor each retry: fixed, still wrong, wrong in a new valid waywhether the gate's message helps or opens an easy exit (our null ids)
Latency per stepwall clock per node from traceswhich step to shrink, parallelise or give a smaller model

Discussion

  1. How many steps should a chain have? One per distinct operation, and no more. Merge two model steps whose combined prompt is still a single operation, and replace any step that can be code.
  2. Should a step see the original input or only the previous output? Passing only structured outputs keeps contexts small but loses information. A reasonable rule: steps that write for a person see the original input; steps that classify or check see only what they need.
  3. Retry or repair? LinkedIn chose a parser for latency. My view: repair what you can predict, retry what you cannot, and count both in the trace, because a rising repair rate is often the first sign that an upstream prompt or model has changed.

4.3 Routing: classify, then dispatch

The example: not every Northwind question needs the expensive model

Northwind's team also runs an internal help desk where staff ask the model quick questions. Most are easy ("what is the plural of criterion?", "what is 15% of 240?"); some are multi-step calculations of the kind that price disputes produce ("3 pens for £2.40 and notebooks at £1.75; what do 9 pens and 4 notebooks cost?"). Sending everything to the strongest model is accurate and expensive; sending everything to the cheapest is cheap and wrong on the hard ones. A router looks at each input and sends it to the handler that suits it.

classifier routercheap-first cascadeinputroutersmall modelor classifierbilling prompttech promptstrong modelunknown: humanone decision up front; each branch has its ownprompt, model and eval setinputcheap modelconfident?yesanswernostrong model"confident" = two cheap samples agree, ascorer passes, or a check on the output
Two routing designs. A classifier router decides up front and sends the input to a specialised prompt, a stronger model or a person (the "unknown" bucket). A cheap-first cascade answers with the cheap model and escalates only when a confidence check fails.

The arithmetic: a router's confusion matrix is its cost model

A router makes two kinds of mistake. It sends a hard input to the cheap handler (accuracy is lost), or an easy input to the expensive one (money is lost). Write hh for the share of hard inputs, ss for the router's recall on hard inputs (the share of hard inputs it escalates) and ee for its false-escalation rate (the share of easy inputs it escalates). Then the share of traffic that reaches the strong model is

strong share=h s+(1−h) e,\text{strong share} = h\,s + (1 - h)\,e ,

and the cost per question is the router's own cost plus the cheap and strong costs weighted by their shares. Assumptions: the two error rates do not depend on the input, and each model's accuracy on easy and hard inputs is fixed. ch4_cost.py tabulates this with illustrative accuracies; the point it makes is that every point of lost recall costs accuracy and every point of false escalation costs money, so you cannot judge a router by "accuracy" alone. You need its confusion matrix.

Our experiment: a cascade and a classifier router

ch4_router.py answers 24 questions with exact answers (a mixed set standing in for the help desk: eight lookups, eight multi-step calculations, eight trickier counting and date problems) in five ways: always-cheap, always-strong, a cascade (two cheap samples; if they agree, accept, otherwise ask the strong model), a classifier router (one cheap call labels the question easy or hard), and an oracle router that knows which questions the cheap model gets right. "Cheap" is gpt-5-mini at minimal reasoning effort; "strong" is the same model at medium effort, so the price per token is identical and the difference is the hidden reasoning tokens the medium setting spends. That keeps the comparison on one verified price, and it means our quality gap is smaller than two different models would show. The five conditions differ only in the routing policy.

Grading is typed and exact. An earlier draft of this book graded by substring, which gives "184" credit for 84; ch4_grade.py instead requires exactly one number and compares it numerically, compares times as minutes, and compares words after explicit normalisation. It prints its own test on adversarial strings:

Terminal output of ch4_grade.py: nineteen adversarial cases, including 184 against gold 84 graded False, 13 against 3 False, It is not 84; the answer is 74 against 84 False because it contains two numbers, The answer is 84 True, 84.0 True, 1,200 g against 1200 True, 14.25 against 14.20 False, 2496 with working shown False, Saturday or Sunday False, Lisbon, Portugal False; 19 of 19 cases behave as specified

python
# code/agents/ch4_router.py (abridged): the two routing policies, computed from the same samples so they differ only in the policy.
def cascade(r):                      # two cheap samples; accept if they agree, otherwise pay for the strong model
    if r['agree']:
        return r['cheap_ok'], [r['u_cheap'], r['u_cheap2']]
    return r['strong_ok'], [r['u_cheap'], r['u_cheap2'], r['u_strong']]

def router(r):                       # one cheap classification call decides before anything is answered
    if r['route'] == 'easy':
        return r['cheap_ok'], [r['u_router'], r['u_cheap']]
    return r['strong_ok'], [r['u_router'], r['u_strong']]

Terminal output of ch4_router.py: a per-question table of 24 rows with the cheap and strong answers, whether two cheap samples agreed and where the router sent each; then the summary: always-cheap 17 of 24, 71 percent, 0.048 dollars per thousand correct answers; always-strong 24 of 24 at 0.194; cascade 21 of 24, 88 percent, at 0.135; classifier router 17 of 24 at 0.125; oracle router 24 of 24 at 0.092; the cascade escalated 5 of 24 and the router sent 1 of 24 to the strong model; the confusion matrix against the oracle shows 7 questions that needed the strong model all sent to the cheap one

Our router experiment: 24 questions, cheap = minimal effort, strong = medium effortaccuracy$ per 1,000 correct answersalways-cheapalways-cheap: 71%71%always-strongalways-strong: 100%100%cascadecascade: 88%88%classifier routerclassifier router: 71%71%oracle routeroracle router: 100%100%: 0.0480.048: 0.1940.194: 0.1350.135: 0.1250.125: 0.0920.092the classifier router against the truthsent to cheapsent to strongcheap suffices161needs strong70green: right decision; orange: wrong.Bottom-left costs accuracy (a hardquestion answered cheaply); top-rightcosts money. Ours sent 1 of 24 questionsto the strong model and caught 0 of the 7that needed it.
Results of the final run of ch4_router.py. The cascade recovered most of the strong model's accuracy at about two-thirds of its cost per correct answer; the classifier router sent only one question to the strong model, caught none of the seven that needed it, and ended with the cheap model's accuracy at more than twice its cost per correct answer.

Three results, each a general point.

The classifier router failed in the way routers usually fail: it was overconfident. Asked "would a small model handle this?", the cheap model said yes to 23 of 24 questions, including all seven it then got wrong. Its recall on the hard questions was zero, so it matched always-cheap's 71% while paying for an extra call on every question. A router's error is invisible in end-to-end accuracy until you build the confusion matrix, and here the matrix is the whole story.

The cascade worked, and its check was weaker than it looked. It escalated 5 of 24 questions and reached 88% (21 of 24) at $0.135 per thousand correct answers, against $0.194 for always-strong. But read its agreement check as a verifier, which ch4_router_stats.py does from the saved results. The check accepted 16 of the 17 right cheap answers (rr = 94%) and also 3 of the 7 wrong ones (ff = 43%): on "what is the sum of the digits of 2 to the power 20?", both cheap samples said 7 (the answer is 31). When a model is confidently wrong, it is wrong twice.

Terminal output of ch4_router_stats.py: p, the chance the cheap answer is right, is 17 of 24, 71 percent; r, the chance two samples agree given a right answer, 16 of 17, 94 percent; f, the false-accept rate, 3 of 7, 43 percent; precision, the chance an accepted answer is right, 16 of 19, 84 percent, matching the Bayes formula; three accepted but wrong answers listed

The false-accept rate is not one minus precision. Precision, the share of accepted answers that are right, depends on how often the cheap model is right in the first place. With pp the share of right cheap answers, Bayes' rule gives

P(right∣accepted)=p rp r+(1−p) f.P(\text{right} \mid \text{accepted}) = \frac{p\,r}{p\,r + (1-p)\,f}.

For our cascade, p=0.71p = 0.71, r=0.94r = 0.94 and f=0.43f = 0.43 give 84%: 16 of the 19 accepted answers were right. The same check on a task where the cheap model is right 30% of the time would give much lower precision. Assumptions: rr and ff are properties of the check that do not change with the mix of questions, which is only roughly true.

The research behind routing

Routing between models of different price became a research topic in 2023, when API prices differed by two orders of magnitude. Lingjiao Chen, Matei Zaharia and James Zou at Stanford posted FrugalGPT in May 2023 and made the cascade the reference design.

Two later papers trained routers that decide before any answer is generated. Microsoft's Hybrid LLM (Ding et al., ICLR 2024) routes between a small and a large model by predicted difficulty and reports "up to 40% fewer calls to the large model, with no drop in response quality". RouteLLM (Ong et al., UC Berkeley, Anyscale and Canva, June 2024) trains routers on human preference data from Chatbot Arena and reports cost reductions of "over 2 times" at matched quality. Both report their gains as curves of quality against the share of calls sent to the strong model, which is the honest way to show a router: a single accuracy number hides the trade.

Research notes: RouteLLM, Hybrid LLM and the FrugalGPT savings table (optional reading)

AWS's prescriptive guidance on agentic patterns (July 2025) lists the same use cases from the practitioner's side: "triaging requests across a variety of tasks", inputs that "must be preprocessed or normalized before entering more specialized workflows", and an agent "acting as a conversational switchboard".

Routing in practice

Classifier router (decide first)Cascade (answer cheaply, then check)
Prosone decision, then a specialised handler with its own prompt, model and eval set; can send "unknown" to a personneeds no labelled routing data to start; the strong model is paid for only on escalations
Consits mistakes are silent unless you build its confusion matrix; a model asked "is this hard?" tends to say noevery input pays for at least one cheap call; a weak check accepts confident wrong answers
When to pick itinputs fall into distinct kinds with different prompts or toolsone kind of task, cheap and strong models of the same family, and a check you have measured
Production failurea new kind of input is forced into the nearest branchthe cheap model drifts and the check keeps accepting

What to measure before you decide

MeasurementHowWhat it tells you
Router confusion matrixlabel a few hundred inputs by outcome (did the cheap handler get it right?)lost accuracy and wasted cost, separately
Check rr and fffor a cascade, how often the check accepts right and wrong cheap answersprecision via Bayes; whether the cascade is safe
Strong-model sharethe share of traffic escalatedcost per thousand requests, and how it moves as traffic changes
Unknown-bucket rateinputs the router sends to a person or a fallbackcoverage; a rising rate is new kinds of input arriving

Discussion

  1. Rule, classifier or model? A rule is free and exact where a field decides; a trained classifier is cheap and measurable; a model handles the long tail. Most production routers are all three in that order.
  2. Can the cheap model be its own router? Ours said yes to 23 of 24 questions. My view: only if you have measured its self-assessment against outcomes; a separately trained scorer, as in FrugalGPT, is usually better.
  3. Where does a router go in a chain? Usually first, so each branch can be a simpler chain. A second router deep in a chain is often a sign the first one is too coarse.

4.4 Parallelisation: split the work, or do it several times

The example: two different reasons to make several calls at once

Before a refund is approved, Northwind wants three checks on the customer's message: does the refund fit the returns policy, is there a fraud signal, is the message abusive? They are independent, so there is no reason to run them one after another. That is sectioning. Separately, the shop's price calculations ("9 pens at 3 for £2.40 and 4 notebooks at £1.75, paid with a £20 note: what change?") are sometimes wrong; asking several times and taking the most common answer might help. That is voting.

sectioning: split the workvoting: same work, n timesinputsection Asection Bsection Cmergeinputsample 1sample 2sample 3votedifferent prompts on different parts; mergeby code (union) or by one model callthe same prompt n times; aggregate by majority,a verifier, or a judge choosing the bestlatency = the slowest branch + the merge; cost = the sum of the branches + the merge. Parallelcalls buy time, never tokens. Voting only helps when samples are right more often than wrong and donot share mistakes.
Sectioning against voting. Sectioning splits a task into independent parts with different prompts and merges them; voting runs the same prompt several times and aggregates. Both run concurrently, so they buy time, not tokens.

The arithmetic: what voting can and cannot do

Suppose each sample is right with probability pp, and when it is wrong it picks one of mm different wrong answers equally often. A plurality vote returns the most frequent answer. It does not need a majority: a right answer given 40% of the time beats two wrong answers given 30% each. ch4_cost.py computes the exact probability that the plurality is right:

right answer ppdistinct wrong answers mmn = 1n = 5n = 11
0.41 (binary)40%32%25%
0.4240%45%50%
0.4540%58%76%
0.6260%77%90%

Assumptions: samples are independent, pp is constant, and wrong answers spread evenly. Two consequences follow. Voting helps when the right answer is the most frequent one, and it helps more when wrong answers scatter; it hurts when one wrong answer is more common than the right one (the binary row). And real samples are not independent: if, with probability ρ\rho, all samples copy one shared misreading of the problem, accuracy becomes ρ p+(1−ρ) plurality\rho\,p + (1-\rho)\,\text{plurality}, which for p=0.4p = 0.4, m=2m = 2, n=5n = 5 falls from 45% at ρ=0\rho = 0 to 42% at ρ=0.6\rho = 0.6. Correlated errors are why voting gains are usually smaller than the formula promises.

With a verifier that can check each candidate (a test, a calculation, a schema), you do not need the most frequent answer, only one that passes. If each sample passes with probability pp, at least one of nn passes with probability 1−(1−p)n1 - (1-p)^n, under the same independence assumption and a verifier that never accepts a wrong answer.

Our experiment: voting, best-of-n with a verifier, and wall-clock time

ch4_parallel.py measures all three with the real model. Voting: ten multi-step word problems with one exact answer (Northwind's price-calculation kind), five samples each, plurality vote over the first 1, 3 and 5. Best-of-n with a verifier: ten "Game of 24" puzzles (combine four numbers with + − × ÷ into 24), chosen because a program can check any candidate exactly; a puzzle counts as solved if any of the first nn candidates verifies. Latency: five samples run one after another, then on five threads. The run cost under a cent.

python
# code/agents/ch4_parallel.py (abridged): the same five calls, one after another and then fanned out on threads.
t0 = time.time()
for _ in range(N):
    llm.ask(P24_PROMPT.format(nums=nums), max_tokens=200)          # sequential: latency is the sum
seq = time.time() - t0
t0 = time.time()
with ThreadPoolExecutor(max_workers=N) as ex:                      # fan-out: latency is the slowest call
    list(ex.map(lambda _: llm.ask(P24_PROMPT.format(nums=nums), max_tokens=200), range(N)))
par = time.time() - t0                                             # the token bill is identical in both cases

Terminal output of ch4_parallel.py: ten word problems with five samples each and the plurality vote at n of 1, 3 and 5, accuracy 40 percent at every n while cost rises to 5 times; ten Game of 24 puzzles with the verified samples marked, solved 10, 30 and 30 percent at n of 1, 3 and 5 against a per-sample pass rate of 18 percent, which would predict 45 and 63 percent if samples were independent; wall-clock time of five samples, 4.6 seconds one by one against 1.0 seconds on five threads, a 4.5 times speedup

Our parallel experiment: accuracy against samples, and wall-clock time0255075100135vote, word problems, n=1: 40vote, word problems, n=3: 40vote, word problems, n=5: 40best-of-n, Game of 24, n=1: 10best-of-n, Game of 24, n=3: 30best-of-n, Game of 24, n=5: 30if independent, n=1: 18if independent, n=3: 45if independent, n=5: 63if independentvote, word problemsbest-of-n, Game of 24samples per task (n)solved, %five samples, mean of 3 runsone by oneone by one: 4.6 s4.6 s5 threads5 threads: 1.0 s1.0 ssame tokens, same bill;4.5x faster when fanned outVoting stayed at 40%: on the problems the model missed, its samples scattered over wrong answersand no majority was right. Best-of-n with a verifier rose from 10% to 30%, far below the 63% thatindependent samples would give: the samples fail on the same puzzles. Cost grew exactly n times inboth.
Results of the final run of ch4_parallel.py. The plurality vote stayed at 40% from one sample to five; best-of-n with a verifier rose from 10% to 30%, well below the 63% that independent samples would give. Fanning five calls out on threads cut wall-clock time from 4.6 to 1.0 seconds for the same tokens.

Voting did not help, and the samples show why. On the four problems the model solved, all five samples agreed (or four of five); on the six it missed, the samples scattered over wrong answers ("3.2", "6.15", "3.25", "5.2", "5.6" for a right answer of 5.80) or agreed on the same wrong one ("80" four times out of five for a right answer of 60). In neither case was the right answer the most frequent. Voting amplifies a model that is usually right on that question; it cannot create an answer the model rarely produces.

Best-of-n with a verifier helped, by less than independence predicts. Across all fifty candidates, 18% verified. If samples were independent, three would solve 45% of puzzles and five would solve 63%; we measured 30% and 30%. The verified candidates clustered on the same three puzzles, and on seven puzzles no candidate verified at all. Samples from one model at one temperature share their blind spots.

Fan-out bought time, not tokens. Five calls took 4.6 seconds one by one and 1.0 second on five threads, a 4.5 times speedup, for the same bill. For sectioning that is the whole point; for voting it means the latency of nn samples is about the latency of one.

The research behind parallel sampling

The research line runs from self-consistency (Chapter 3) through two 2024 papers that scaled sampling far further. Bradley Brown and colleagues at Stanford, Oxford and Google DeepMind asked in "Large Language Monkeys" (July 2024) what happens with hundreds or thousands of samples per problem.

Two other papers fill in the picture. "More Agents Is All You Need" (Li et al., Tencent, TMLR 2024) samples up to forty answers and votes, and finds gains that grow with task difficulty: with Llama2-13B, GSM8K accuracy went from 0.35 with one sample to 0.59 with forty, above Llama2-70B's single sample (0.54). Universal Self-Consistency (Chen et al., Google, November 2023) handles free-form answers, where exact voting is impossible, by asking a model to pick "the most consistent response based on majority consensus"; it matched exact-match voting on maths and execution-based voting on SQL without running the code.

Research notes: More Agents, Universal Self-Consistency and the Monkeys abstract (optional reading)

Parallelisation in practice

SectioningVoting or best-of-n
Proslatency of the slowest section, not the sum; each section has a focused prompt and its own evalcan lift accuracy when samples vary and the right answer is common, or when a verifier exists
Conssections must really be independent, or the merge has to reconcile them; cost is the sumnn times the tokens; correlated errors cut the gain; a weak aggregator picks wrong
When to pick itseparate checks or separate parts of a documentexact answers with high disagreement between samples, or any task with a cheap verifier
Production failuretwo sections disagree and the merge silently picks onea vote that is confidently wrong because every sample made the same mistake

What to measure before you decide

MeasurementHowWhat it tells you
Section independencedoes any section's output change another's?whether you can fan out at all
Agreement ratesample the same input five timesif samples agree, one is enough; if the right answer is rare, voting cannot help
Per-sample pass rate with a verifierverify every candidate on a labelled setthe ceiling for best-of-n, and how far correlation pulls you below 1−(1−p)n1-(1-p)^n
Wall-clock time, sequential against concurrenttime both on real trafficthe speedup your provider's rate limits actually allow

Discussion

  1. Is sectioning just a chain with the arrows removed? Only when the sections do not depend on each other. If the fraud check needs the policy check's result, you have a chain, and running it in parallel creates a race.
  2. When is voting worth five times the tokens? When samples disagree often and the right answer is usually the most common one. Measure both on a hundred inputs first; our run would have saved 80% of its voting cost by checking agreement first.
  3. Different prompts or the same prompt? Varying the prompt or the model across samples reduces correlated errors, at the cost of comparability. Uber's code review (Section 4.8) uses different specialised assistants, which is sectioning, not voting.

4.5 Orchestrator-workers: when the split cannot be written in advance

The example: a question nobody can decompose in advance

Northwind's head of operations asks: "Delivery complaints rose by about forty per cent last quarter. Why, and what should we change?" No fixed chain answers this. The work depends on what the first look finds: if complaints cluster in two regions, someone has to read those regions' courier notices; if they cluster on one product, someone has to read its returns data. The number and kind of subtasks are decided by the input and by early findings. That is the case for an orchestrator: a model that plans the split, hands each part to a worker, and combines what comes back.

Orchestrator-workers: the split is decided per input, at run timetaskorchestratorplans, writes onebrief per workerworker 1own context, own toolsworker 2own context, own toolsworker kown context, own toolssynthesiseanswernot enough yet: plan again, spawn more workersthe difference from sectioning: there the sections are written in code; here the model decides howmany and which. what goes wrong: vague briefs (workers duplicate each other), too many workers foran easy task, a synthesis that drops or contradicts findings. What it costs: the sum of everyworker's whole loop; latency is the slowest worker.
Orchestrator-workers. The orchestrator plans and writes a brief per worker; workers run with their own contexts and tools, usually in parallel; a synthesis step combines their results, and the orchestrator can plan again and spawn more workers. Typical failures are vague briefs (duplicated work), too many workers for an easy task, and a synthesis that drops or contradicts findings.

What the production evidence says

The clearest public account is Anthropic's description of the multi-agent system behind Claude's Research feature, published in June 2025.

The same post lists what went wrong early, in terms every orchestrator builder will recognise: agents "spawning 50 subagents for simple queries", "scouring the web endlessly for nonexistent sources", and subagents that "duplicated work" because the lead agent's briefs were short ("research the semiconductor shortage"). The fixes were in the orchestrator's prompt: briefs with "an objective, an output format, guidance on the tools and sources to use, and clear task boundaries", and explicit effort rules ("simple fact-finding requires just 1 agent with 3-10 tool calls ... complex research might use more than 10 subagents"). Running 3 to 5 subagents in parallel, each making several tool calls in parallel, "cut research time by up to 90% for complex queries".

The cost arithmetic, with its assumptions

ch4_cost.py compares one agent doing kk subtasks in sequence with an orchestrator that writes kk briefs, runs kk workers concurrently and synthesises. With a 1,500-token prompt, 300-token results and a 400-token brief, the orchestrator sends 20,600 input tokens at k=8k = 8 against 29,700 for the single agent, and finishes in 25 seconds against 72.

One agent doing k subtasks in turn, against an orchestrator with k workersinput tokens per tasklatency per task (s)k=2, one agentk=2, one agent: 5,8505,850k=2, orchestratork=2, orchestrator: 7,4007,400k=4, one agentk=4, one agent: 12,00012,000k=4, orchestratork=4, orchestrator: 11,80011,800k=8, one agentk=8, one agent: 29,70029,700k=8, orchestratork=8, orchestrator: 20,60020,600: 2424: 2525: 4040: 2525: 7272: 2525Planning numbers of ch4_cost.py (a worker gets a 400-token brief plus the 1,500-token prompt).With fresh worker contexts the tokens are similar to a loop that re-reads its history, and thetime is the slowest worker. Anthropic's 15x multiplier comes from each worker running a full tool-using loop of its own, which this arithmetic leaves out.
Input tokens and latency per task for one agent doing k subtasks in sequence and for an orchestrator with k concurrent workers, from ch4_cost.py. Under these assumptions the tokens are similar and the orchestrator is faster by the width of the fan-out; real workers that run their own tool loops multiply the token count.

Assumptions, and why they matter: the single agent replays its whole growing history at every step (no compaction, no prompt caching, both of which cut its cost a lot); each worker makes exactly one call; workers run concurrently; the orchestrator does not re-plan. Real workers usually run a tool-using loop of their own, which is where Anthropic's "about 15×" comes from. So the arithmetic says something narrower than "orchestrators are cheap": orchestration itself adds little; what costs is many workers each doing real work, and what you buy is time and breadth.

The research behind orchestrators

The pattern appeared in research in 2023. HuggingGPT (Shen et al., Zhejiang University and Microsoft Research Asia, March 2023) used a language model as a controller over expert models from the Hugging Face hub in four stages: task planning, model selection, task execution and response generation. AutoGen (Wu et al., Microsoft Research and three universities, August 2023) made multi-agent conversation a programming framework. Magentic-One (Fourney et al., Microsoft Research, November 2024) is the most carefully engineered of the three, and its orchestrator keeps two explicit records.

Research notes: HuggingGPT, Magentic-One's ledgers and results, AutoGen, and Anthropic's prompting rules (optional reading)

Orchestrator-workers in practice

Fixed sectioningOrchestrator-workers
Prospredictable cost and latency; each section testablehandles tasks whose split depends on the input or on early findings; breadth in parallel
Conscannot adapt the splitcost is the sum of every worker's loop (Anthropic reports about 15 times a chat); plans and briefs can be wrong; harder to test
When to pick itthe parts are known in advanceopen-ended, breadth-first, high-value tasks with independent sub-questions
Production failurea part nobody planned forduplicated or contradictory workers, runaway spawning, a synthesis that drops findings

What to measure before you decide

MeasurementHowWhat it tells you
Plan qualitygrade a sample of plans before executionwhether the orchestrator decomposes well, separately from execution
Worker overlapcompare workers' queries or sources per taskduplicated work from vague briefs
Workers and tokens per taskfrom traces, against task difficultywhether effort scales with the question (Anthropic's early failure)
Synthesis fidelitycheck that each worker's key finding appears in the answerlost or contradicted findings
Value per taskwhat a good answer is worth against its token costwhether the pattern pays at all

Discussion

  1. Is orchestrator-workers a workflow or an agent? The orchestrator decides the split, so it is the most agent-like of the five. In practice the outer frame (plan, fan out, synthesise, stop) is code and the inside is the model's, which is where the testing burden lands.
  2. Should workers share context? Isolation keeps contexts small and stops one worker's error spreading; sharing lets workers build on each other. My view: isolate by default and pass findings explicitly through the orchestrator, because that path is traceable.
  3. When is the 15-times cost worth it? When a wrong or slow answer is expensive and the question is breadth-first. A monthly operations question at Northwind qualifies; answering each support ticket does not.

4.6 Evaluator-optimiser: generate, check, refine, stop

The example: rewriting Northwind's weekly update

Every week Northwind's team leads write paragraphs about their work, and the internal-communications editor wants them rewritten for the all-staff email under strict rules: 55 to 75 words, exactly four sentences, no sentence over 20 words, three named keywords, and none of ten banned words ("very", "basically", "leverage" and so on). A single call often misses one rule. The evaluator-optimiser pattern adds a second role: something that checks the draft and sends back what is wrong, in a loop that stops by a rule.

Evaluator-optimiser: generate, evaluate, feed back, stop by a ruletaskgeneratordraftevaluator1. code checks first2. a judge modelstop?passshipfail: the exact failures go back to the generator ("C3: a sentence of 23 words, limit 20")stop rules, in the order to apply thempassesevery check passes, or the score clears a thresholdmax roundstwo to four; most of the gain comes by the thirdno progressthe same failure twice in a row: stop, do not loopthenhand the best draft so far to a person, with its failures
The evaluator-optimiser loop. The generator writes a draft; the evaluator runs code checks first and a judge model second; a stop rule decides; on failure the exact failures go back to the generator. Below, the stop rules in the order to apply them.

The arithmetic of the stop, with an imperfect evaluator

If each round passes with probability pp and the evaluator is perfect, the chance of passing within kk rounds is 1−(1−p)k1-(1-p)^k and the cost per accepted draft is one round's cost divided by pp, whatever kk is (Chapter 3 derived this). Assumptions: rounds are independent with the same pp, each round costs the same, and the evaluator never errs. Feedback-driven rounds break the first assumption (feedback usually raises pp after the first round), so this is a floor on the benefit, not a forecast.

The evaluator is never perfect. Write rr for the chance it accepts a good draft and ff for the chance it accepts a bad one (its false-accept rate). The share of drafts the loop stops on that are actually good is, by Bayes' rule,

P(good∣stopped)=p rp r+(1−p) f,P(\text{good} \mid \text{stopped}) = \frac{p\,r}{p\,r + (1-p)\,f},

the same formula as the cascade's check in Section 4.3. With p=0.9p = 0.9, r=1r = 1 and f=0.2f = 0.2 it gives 97.8%; with p=0.3p = 0.3 and the same evaluator, 68%. The harder the task, the more the evaluator's false-accept rate matters. And a low rr (rejecting good drafts) has its own cost: more rounds, and rewrites of drafts that were already fine.

Terminal output of part 5 of ch4_cost.py: with a perfect critic, the chance of passing within k rounds and the expected cost, where cost per accepted draft equals one round's cost divided by p for every k; then the imperfect critic table: p good 0.9, r 1.00, f 0.20 gives 92 percent accepted and 97.8 percent precision; p 0.5 gives 83.3 percent; p 0.3 gives 68.2 percent; p 0.5 with r 0.9 and f 0.1 gives 90.0 percent

Our experiment: a code checker against a judge model

ch4_evaluator.py rewrites eight source paragraphs under the five rules, in two loops from the same start, each allowed up to four rounds. In the programmatic loop a Python checker lists the rules each draft breaks, with measured values ("C3: a sentence of 24 words, limit 20"), and the loop stops when the list is empty. In the judge loop a separate model call marks each rule pass or fail, its verdict is fed back, and the loop stops when it says everything passes. Every draft in both loops is also scored by the other evaluator, so we can count how often the judge and the code disagree. Because the programmatic loop is graded by the same code that gives it feedback, we add a held-out check that no loop ever sees: two key facts per source paragraph (for example "staging" and "audit log" in the incident report) must survive the rewrite. The run cost about two cents.

python
# code/agents/ch4_evaluator.py (abridged): one loop; arm decides who the critic is. The held-out fact check is applied once, at the end.
for rnd in range(1, MAX_ROUNDS + 1):
    text, ug = llm.ask(msgs, max_tokens=400)                  # generator
    viol = check(text, keywords)                              # programmatic checker: exact counts, never wrong about counting
    jfails, notes, uj = judge(text, keywords)                 # judge model: pass/fail per constraint
    critic_fails = viol if arm == 'programmatic' else jfails  # who decides whether to stop
    if not critic_fails:
        break
    feedback = viol if arm == 'programmatic' else [f'{c}: the reviewer marked this constraint as failed. {notes}' for c in jfails]
    msgs += [{'role': 'assistant', 'content': text}, {'role': 'user', 'content': REGEN.format(viol='\n- '.join(feedback))}]

Terminal output of ch4_evaluator.py: per-paragraph trails for both loops; then over 40 drafts the judge agreed with the checker 40 percent of the time, with 2 false passes and 22 false fails; agreement by constraint 32 percent on the word count, 90 on four sentences, 62 on the sentence length limit, 90 on keywords and 80 on banned words; summary: the programmatic loop stopped on 7 of 8, all 7 truly passing, mean 1.88 rounds; the judge loop stopped on only 4 of 8 after a mean of 3.12 rounds, although 7 of its final drafts passed every constraint; held-out facts kept 8 of 8 and 7 of 8; cost per round 0.0002 against 0.0005 dollars

Our evaluator experiment: a judge model against a programmatic checkerjudge agrees with the checker, per constraintfinal drafts (8 paragraphs)C1 55-75 wordsC1 55-75 words: 32%32%C2 four sentencesC2 four sentences: 90%90%C3 sentence <= 20C3 sentence <= 20: 62%62%C4 keywordsC4 keywords: 90%90%C5 banned wordsC5 banned words: 80%80%code: loop stoppedcode: loop stopped: 7/87/8code: truly passcode: truly pass: 7/87/8judge: loop stoppedjudge: loop stopped: 4/84/8judge: truly passjudge: truly pass: 7/87/8Over 40 drafts the judge agreed with the checker 40% of the time: 2 false passes and 22 falsefails. It was worst at counting words (C1, 32%). In this run its errors were mostly false fails,so the judge loop ran 3.1 rounds on average against 1.9 and stopped on only 4 of 8 paragraphs,although 7 of its final drafts met every constraint. A held-out fact check, never fed back, passed8/8 and 7/8.
Results of the final run of ch4_evaluator.py. The judge agreed with the code checker on only 32% of word-count verdicts. In this run its errors were mostly false fails, so its loop ran 3.1 rounds on average against 1.9 and cost more than twice as much per round, and rewriting drafts that already passed cost one paragraph a key fact.

The judge cannot count. Over 40 drafts it agreed with the checker's overall verdict 40% of the time, with 2 false passes and 22 false fails. It was worst on the word count (32% agreement) and on the sentence-length rule (62%), the two rules that require counting, and best on keywords (90%), which require only looking.

False fails are not harmless. In this run the judge mostly rejected drafts that were fine. Its loop averaged 3.1 rounds against 1.9, five of its eight paragraphs hit the four-round cap, and with the judge's own call in every round it cost $0.0005 per round against $0.0002. Worse, rewriting a passing draft can damage it: on paragraph 5 the first draft met every rule, the judge rejected it three times, and the final draft had dropped one of the held-out facts (the hourly queries that drive the warehouse costs). In a development run the judge's errors leaned the other way, towards false passes; the direction varies, the unreliability does not.

A false pass ships a broken draft. On paragraph 1 the judge accepted a 52-word draft against a 55-word minimum after one round. The programmatic loop, whose checker cannot miscount, stopped on 7 of 8 paragraphs, and all 7 truly passed and kept their held-out facts. Its one failure was a paragraph about app crashes where the model kept writing one 24-word sentence across four rounds, the "no progress" case for which a stop rule should hand the draft to a person.

The research behind judges

The reference study of model judges is Lianmin Zheng and colleagues' "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (the LMSYS group at Berkeley and others, June 2023). It found strong judges agreeing with human preferences about as often as humans agree with each other, "over 80% agreement", and it catalogued the biases.

G-Eval (Liu et al., Microsoft, March 2023) showed how to make a judge more consistent: have the model write evaluation steps from the criteria, fill in a form, and weight the score by the probabilities of each grade. Its Spearman correlation with human ratings of summaries was 0.514 with GPT-4, ahead of every earlier automatic metric; the authors also warn that model evaluators may prefer model-written text. Chapter 3's caution applies here too: Huang et al. (2023) found that asking a model to correct its own reasoning without outside feedback often made it worse. A judge is outside feedback only to the extent that it sees something the generator did not, or is better than the generator at the specific check.

Research notes: the judge paper's abstract, G-Eval, and Anthropic's conditions for the pattern (optional reading)

Evaluator-optimiser in practice

Code evaluator (tests, schema, counts)Model judge
Prosexact, free, fast; never miscounts; specific feedbackjudges what code cannot: tone, helpfulness, faithfulness to a source
Consonly checks what you can write downfalse passes and false fails; position and verbosity bias; one extra call per round
When to pick itany rule you can express in codequalities with no programmatic test, after measuring its rr and ff
Production failurea check that tests the wrong thing, passed reliablya judge that drifts with a model update and nobody re-measures it

What to measure before you decide

MeasurementHowWhat it tells you
Judge rr and ffrun the judge on a few hundred labelled draftsthe precision of the stop, via Bayes
Position and verbosity sensitivityswap order; pad an answer without adding contentwhether the judge's verdicts survive changes that should not matter
Rounds to pass, and the no-progress ratefrom traceswhere to cap rounds; how often to hand off to a person
Damage from extra roundsheld-out checks on drafts that passed earlywhether false fails are degrading good work
Cost per round and per accepted draftAPI usage per roundthe price of each extra round

Discussion

  1. Is a judge outside feedback? Partly. It sees the draft fresh, without the generator's reasoning, but it shares the model's blind spots if it is the same model. Using a different model as judge reduces shared blind spots; it does not remove the need to measure the judge.
  2. What should the loop optimise? Whatever the evaluator checks, exactly as our gated chain did. If the evaluator is narrow, add held-out checks the loop never sees, as we did with the key facts.
  3. How many rounds? Our programmatic loop passed most paragraphs by round two; most of the judge loop's extra rounds were spent rewriting drafts that already passed. A cap of two to four rounds, with a no-progress stop, is a reasonable start, then adjust from traces.

4.7 Composing patterns: the workflow as a state machine

The example: Northwind's support workflow, all together

Real systems combine the patterns. Northwind's support workflow, built in full in ch4_workflow.py, puts a router in front (code rules first, a model call only when the rules do not decide), a gated chain behind it (extract, then draft, each schema-validated with one retry), a policy check in code, a human approval step for refunds over £50, and an idempotent send. Around those nodes sits the plumbing every production workflow needs: typed state, budgets, checkpoints, resume, and a trace. A larger version could hang an orchestrator off one branch (for "why are complaints up?" questions) and an evaluator loop before the send; the plumbing would not change.

Patterns compose: router in front, chain or orchestrator behind, evaluator lastinputrouterstep 1step 2step 3known request: a gated chainorchestratorworkers in parallelopen request: orchestrator-workersevaluatorhuman approvesinterrupt: pauseand waitdurable execution: what happens when the process dies at step 3routerstep 1step 2step 3resumestep 3evaluatecrashload the last checkpointSaved after each node (squares): the state and a trace id. Resuming re-runs only step 3, so step 3must be idempotent: running it twice must not send two emails or charge twice. Give each side effectan idempotency key; log one span per node.
Patterns compose. A router sends known requests to a gated chain and open-ended ones to an orchestrator with parallel workers; an evaluator checks the result and a human-approval interrupt pauses the run. Below, durable execution: state is checkpointed after each node, a crash resumes from the last checkpoint, and the re-run step must be idempotent so that it cannot send or charge twice.

The complete program

The core of the runner is a loop over the current node, with a checkpoint after every node:

python
# code/agents/ch4_workflow.py (abridged): the runner. Every node updates the typed State and names the next node.
def run(st, crash_after=None):
    try:
        while st.node not in ('done', 'needs_human', 'awaiting_approval'):
            t0 = time.time()
            if st.node == 'extract':
                obj, use, retried = extract(st)                      # schema-validated (pydantic), one retry with the errors
                if obj is None:
                    span(st, 'extract', 'gate failed twice -> person', t0, use); st.node = 'needs_human'
                else:
                    st.facts = obj.model_dump(); span(st, 'extract', 'ok', t0, use); st.node = 'draft'
            elif st.node == 'review':
                amt = st.facts.get('refund_amount_gbp') or 0
                if amt > REFUND_LIMIT and st.approved is None:
                    span(st, 'review', 'PAUSE for approval', t0); st.node = 'awaiting_approval'   # human-in-the-loop interrupt
                else:
                    span(st, 'review', 'policy ok', t0); st.node = 'send'
            ...                                                      # route, faq, draft and send follow the same shape
            checkpoint(st)                                           # state saved after every node: resume starts here
    except BudgetExceeded as e:                                      # at most 6 model calls and 8,000 tokens per ticket
        span(st, st.node, f'STOP: {e}', time.time()); st.node = 'needs_human'; checkpoint(st)
    return st

The demonstration runs six tickets, kills the process on purpose after ticket T3's draft is checkpointed, resumes every unfinished ticket from its checkpoint, has a person approve the paused refund, and then replays a send to show that it cannot happen twice. It made 12 model calls and cost $0.0015.

Terminal output of ch4_workflow.py: one trace line per node with a trace id, the node, the outcome, tokens and milliseconds; T1 is routed by rule, extracted, drafted, reviewed and sent; the run is killed after T3's draft is checkpointed; T4 is routed by the model to the FAQ workflow; T5's refund of 649 pounds pauses for approval; T6, in Spanish, is routed by the model; on resume T3 continues at the review node without re-running extract or draft; a person approves T5, which is then sent; replaying the send for T1 reports already sent, idempotency key seen; the final table shows every ticket done with two or three calls each, five replies in the outbox and 29 spans in the trace

Four behaviours in that output are the reasons the plumbing exists. Resume skips finished work: T3 continued at review with its two earlier calls already counted, and extract and draft were not run again. The interrupt waits: T5's £649 refund stopped at awaiting_approval with its state on disk, and continued from there after approval. The side effect is idempotent: replaying T1's send found its key in the outbox and did nothing. The router used the model only when needed: four tickets were routed by keyword rules at no cost; the Spanish ticket (T6) and the shipping question (T4) fell through the rules and cost one small call each.

The full program: code/agents/ch4_workflow.py (about 260 lines; optional reading)
python
"""Chapter 4, Section 4.7: one complete, runnable composed workflow for Northwind Home's support tickets, with the real API
(gpt-5-mini, reasoning_effort=minimal). Everything the chapter recommends, in about 250 lines:
  typed state (a dataclass) and typed step outputs (pydantic schemas, validated in code);
  a router that is code first and asks the model only when the rules do not decide;
  a gated chain: extract -> draft, each step schema-validated, retried once with the exact errors, then handed to a person;
  a policy check in code (word limit, banned phrases, order id quoted) after the draft;
  a human-in-the-loop interrupt: refunds over GBP 50 pause the run until someone approves;
  an idempotent side effect: replies go to an outbox keyed by an idempotency key, so a resumed run cannot send twice;
  budgets per ticket (model calls, tokens) that stop the run instead of letting it spin;
  a checkpoint after every node, and resume from the last checkpoint after a crash;
  a trace: one JSON line per node (trace id, node, outcome, tokens, milliseconds) in results/ch4_workflow_trace.jsonl.
The demo processes six tickets, injects one crash after the draft step of ticket T3, resumes it, and approves the paused refund."""
import hashlib, json, os, re, time, uuid
from dataclasses import dataclass, field, asdict
from typing import Literal, Optional
from pydantic import BaseModel, Field, ValidationError
from common import Log, save, RESULTS
import ch4_llm as llm

log = Log('ch4_workflow')
llm.set_log(log)
CKPT = os.path.join(RESULTS, 'ch4_workflow_ckpt')
OUTBOX = os.path.join(RESULTS, 'ch4_workflow_outbox.json')
TRACE = os.path.join(RESULTS, 'ch4_workflow_trace.jsonl')
MAX_CALLS, MAX_TOKENS, REFUND_LIMIT = 6, 8000, 50.0
BANNED = ['24 hours', 'immediately', 'guarantee', 'today', 'right away', 'as soon as possible', 'shortly']


# ---------------------------------------------------------------- typed step outputs (the schemas the gates enforce)
class Facts(BaseModel):
    issue_type: Literal['billing', 'delivery', 'damaged_or_faulty', 'account', 'other']
    order_id: Optional[str] = Field(default=None, pattern=r'^ORD-\d{5}$')
    refund_amount_gbp: Optional[float] = Field(default=None, ge=0)


class Draft(BaseModel):
    reply: str = Field(min_length=20)


# ---------------------------------------------------------------- typed workflow state
@dataclass
class State:
    ticket_id: str
    text: str
    trace_id: str = field(default_factory=lambda: uuid.uuid4().hex[:8])
    node: str = 'route'                 # the next node to run; 'done' and 'needs_human' and 'awaiting_approval' are terminal or paused
    route: Optional[str] = None
    facts: Optional[dict] = None
    reply: Optional[str] = None
    approved: Optional[bool] = None
    calls: int = 0
    tokens: int = 0
    notes: list = field(default_factory=list)


class BudgetExceeded(Exception):
    pass


class InjectedCrash(Exception):
    pass


# ---------------------------------------------------------------- infrastructure: checkpoints, trace, model calls with budgets
def checkpoint(st):
    os.makedirs(CKPT, exist_ok=True)
    json.dump(asdict(st), open(os.path.join(CKPT, st.ticket_id + '.json'), 'w'), indent=1)


def load(ticket_id):
    p = os.path.join(CKPT, ticket_id + '.json')
    return State(**json.load(open(p))) if os.path.exists(p) else None


def span(st, node, outcome, t0, use=None):
    rec = {'trace_id': st.trace_id, 'ticket': st.ticket_id, 'node': node, 'outcome': outcome, 'ms': round((time.time() - t0) * 1000),
           'tokens': (use or {}).get('input', 0) + (use or {}).get('output', 0)}
    open(TRACE, 'a').write(json.dumps(rec) + '\n')
    log(f'  [{st.trace_id}] {st.ticket_id} {node:<9} {outcome:<34} {rec["tokens"]:>5} tok {rec["ms"]:>6} ms')


def call(st, prompt):
    if st.calls >= MAX_CALLS or st.tokens >= MAX_TOKENS:
        raise BudgetExceeded(f'budget: {st.calls} calls, {st.tokens} tokens')
    text, use = llm.ask(prompt, max_tokens=600)
    st.calls += 1
    st.tokens += use['input'] + use['output']
    return text, use


def validated(st, prompt, model_cls, extra_check=None):
    """A gated step: call, parse, validate against the schema (and an optional extra check), retry once with the exact errors."""
    total = {'input': 0, 'output': 0}
    for attempt in range(2):
        text, use = call(st, prompt)
        for k in total:
            total[k] += use[k]
        try:
            obj = model_cls.model_validate(llm.parse_json(text) or {})
            problems = extra_check(obj) if extra_check else []
        except ValidationError as e:
            obj, problems = None, [f'{".".join(str(x) for x in err["loc"])}: {err["msg"]}' for err in e.errors()]
        if not problems:
            return obj, total, attempt
        prompt = prompt + '\n\nYour previous answer failed these checks:\n- ' + '\n- '.join(problems) + '\nReply with a corrected JSON object only.'
    return None, total, 1


# ---------------------------------------------------------------- the nodes
def route(st):
    t = st.text.lower()
    if re.search(r'\b(refund|charged|invoice|broken|cracked|damaged|parcel|delivery|order)\b', t):     # code decides when it can
        return 'support', None
    text, use = call(st, 'Is this message a support request about an order or account (answer "support") or a general question '
                         f'about products, shipping zones or opening hours (answer "faq")? Reply with one word.\n\nMessage: {st.text}')
    return ('faq' if 'faq' in text.lower() else 'support'), use


def extract(st):
    prompt = ('Extract facts from this support ticket. Reply with ONE JSON object with keys: "issue_type" (one of billing, delivery, '
              'damaged_or_faulty, account, other), "order_id" (the order number as ORD-NNNNN, five digits, or null), '
              f'"refund_amount_gbp" (a number if the customer asks for money back and states an amount, else null).\n\nTicket: {st.text}')
    return validated(st, prompt, Facts)


def policy_problems(reply, order_id):
    p = []
    if len(reply.split()) > 80:
        p.append(f'the reply has {len(reply.split())} words; the limit is 80')
    hits = [b for b in BANNED if b in reply.lower()]
    if hits:
        p.append(f'the reply promises timing: {hits}; remove these phrases')
    if order_id and order_id not in reply:
        p.append(f'the reply must quote the order id {order_id}')
    return p


def draft(st):
    f = st.facts
    prompt = ('Write a short, polite reply (at most 80 words) to this customer, for Northwind Home. Do not promise any timing. '
              + (f'Quote the order id {f["order_id"]}. ' if f.get('order_id') else '')
              + ('Say that the refund request has been passed for approval. ' if f.get('refund_amount_gbp') else '')
              + 'Reply with ONE JSON object: ' + json.dumps({'reply': '...'}) + f'\n\nTicket: {st.text}\nFacts: {json.dumps(f)}')
    return validated(st, prompt, Draft, lambda d: policy_problems(d.reply, f.get('order_id')))


def send(st):
    """The side effect. Idempotent: the key is derived from the ticket and the reply, and the outbox ignores a key it has seen."""
    key = f'{st.ticket_id}:' + hashlib.sha256(st.reply.encode()).hexdigest()[:12]   # stable across processes, unlike hash()
    box = json.load(open(OUTBOX)) if os.path.exists(OUTBOX) else {}
    if key in box:
        return 'already sent (idempotency key seen)'
    box[key] = {'ticket': st.ticket_id, 'reply': st.reply, 'at': time.strftime('%H:%M:%S')}
    json.dump(box, open(OUTBOX, 'w'), indent=1)
    return 'sent'


# ---------------------------------------------------------------- the graph runner
def run(st, crash_after=None):
    """Run from st.node until the workflow finishes, pauses or fails; checkpoint after every node."""
    try:
        while st.node not in ('done', 'needs_human', 'awaiting_approval'):
            t0 = time.time()
            if st.node == 'route':
                st.route, use = route(st)
                span(st, 'route', f'-> {st.route}' + (' (model)' if use else ' (rule)'), t0, use)
                st.node = 'extract' if st.route == 'support' else 'faq'
            elif st.node == 'faq':
                st.notes.append('sent to the FAQ answerer (out of scope for this demo)')
                span(st, 'faq', 'handed to the FAQ workflow', t0)
                st.node = 'done'
            elif st.node == 'extract':
                obj, use, retried = extract(st)
                if obj is None:
                    span(st, 'extract', 'gate failed twice -> person', t0, use); st.node = 'needs_human'
                else:
                    st.facts = obj.model_dump()
                    span(st, 'extract', f'ok{" after retry" if retried else ""} {st.facts["issue_type"]} {st.facts["order_id"]}', t0, use)
                    st.node = 'draft'
            elif st.node == 'draft':
                obj, use, retried = draft(st)
                if obj is None:
                    span(st, 'draft', 'gate failed twice -> person', t0, use); st.node = 'needs_human'
                else:
                    st.reply = obj.reply
                    span(st, 'draft', f'ok{" after retry" if retried else ""} ({len(st.reply.split())} words)', t0, use)
                    st.node = 'review'
            elif st.node == 'review':
                amt = st.facts.get('refund_amount_gbp') or 0
                if amt > REFUND_LIMIT and st.approved is None:
                    span(st, 'review', f'refund GBP {amt:.2f} > {REFUND_LIMIT:.0f}: PAUSE', t0); st.node = 'awaiting_approval'
                elif st.approved is False:
                    span(st, 'review', 'refund rejected by a person', t0); st.node = 'needs_human'
                else:
                    span(st, 'review', 'policy ok' + (' (approved)' if st.approved else ''), t0); st.node = 'send'
            elif st.node == 'send':
                span(st, 'send', send(st), t0); st.node = 'done'
            checkpoint(st)
            if crash_after and st.node == crash_after[1] and st.ticket_id == crash_after[0]:
                raise InjectedCrash(f'process killed after checkpointing, before running {st.node}')
    except BudgetExceeded as e:
        span(st, st.node, f'STOP: {e}', time.time()); st.node = 'needs_human'; checkpoint(st)
    return st


TICKETS = [('T1', 'Order #48213 arrived with the screen cracked. Please send a replacement.'),
           ('T2', 'I was charged twice for ORD-10988, please refund the extra GBP 39.99.'),
           ('T3', 'Where is my parcel? Tracking has not moved in 6 days. Order 77120.'),
           ('T4', 'Do you ship to the Isle of Man?'),
           ('T5', 'The sofa from order 61234 has a torn seam. I want my money back, GBP 649.'),
           ('T6', 'Mi pedido 33310 llego roto.')]

if __name__ == '__main__':
    import shutil
    for p in [CKPT]:
        shutil.rmtree(p, ignore_errors=True)
    for p in [OUTBOX, TRACE]:
        if os.path.exists(p):
            os.remove(p)
    log(f'== A composed workflow for Northwind Home support ({llm.MODEL}, minimal effort): router -> gated chain -> policy -> approval -> send ==')
    log(f'budgets per ticket: {MAX_CALLS} model calls, {MAX_TOKENS} tokens; refunds over GBP {REFUND_LIMIT:.0f} need a person')
    log('')
    log('-- first run (a crash is injected in T3 after the draft is checkpointed) --')
    for tid, text in TICKETS:
        st = State(tid, text)
        try:
            run(st, crash_after=('T3', 'review'))
        except InjectedCrash as e:
            log(f'  !! {tid}: {e}')
    log('')
    log('-- resume: load every checkpoint that is not finished and continue from its saved node --')
    for tid, _ in TICKETS:
        st = load(tid)
        if st.node not in ('done', 'needs_human', 'awaiting_approval'):
            log(f'  {tid}: resuming at node "{st.node}" (calls so far {st.calls}; extract and draft are NOT re-run)')
            run(st)
    log('')
    log('-- a person approves the paused refund, and the run continues from the checkpoint --')
    for tid, _ in TICKETS:
        st = load(tid)
        if st.node == 'awaiting_approval':
            st.approved = True
            st.node = 'review'
            log(f'  {tid}: approved by a person')
            run(st)
    log('')
    log('-- replaying "send" for T1 to show idempotency --')
    st = load('T1'); t0 = time.time(); span(st, 'send', send(st), t0)
    log('')
    final = [load(t) for t, _ in TICKETS]
    log(f'{"ticket":<7} {"route":<8} {"final node":<12} {"calls":>5} {"tokens":>7}  facts')
    for st in final:
        f = st.facts or {}
        log(f'{st.ticket_id:<7} {str(st.route):<8} {st.node:<12} {st.calls:>5} {st.tokens:>7}  {f.get("issue_type", "-")} {f.get("order_id", "-")} '
            f'{("refund " + str(f.get("refund_amount_gbp"))) if f.get("refund_amount_gbp") else ""}')
    box = json.load(open(OUTBOX))
    spans = [json.loads(l) for l in open(TRACE)]
    U = llm.snapshot()
    log(f'outbox: {len(box)} replies sent for {sum(1 for s in final if s.node == "done" and s.route == "support")} finished support tickets; '
        f'trace: {len(spans)} spans in results/ch4_workflow_trace.jsonl')
    log(f'total spend: ${llm.cost(U):.4f} ({U["calls"]} calls); stub used: ' + ('YES' if llm.stub_used() else 'no'))
    save('ch4_workflow', {'states': [asdict(s) for s in final], 'spans': spans, 'outbox': box, 'usage': U, 'stub': llm.stub_used()})
Research notes: LangGraph's interrupt (optional reading)

Observability to design in now

Every node in the program writes one line to a trace file: the trace id of the run, the ticket, the node, the outcome, the tokens and the milliseconds. That is the minimum. With it you can answer the questions this chapter has asked of every pattern: which step fails first, how often each gate retries, how many tickets the router sends to the model, how long approvals wait, what each ticket costs. Chapter 8 builds this into proper tracing with spans, parent ids and dashboards; the decision to make now is that every node emits a record with the same trace id.

Composition in practice

PlumbingWhat it preventsHow the program does it
Typed state and schemasa step reading a field that is missing or malformeda dataclass for state; pydantic models for step outputs
Gates with one retryformat failures flowing downstreamvalidated(): parse, validate, retry once with the errors
Budgets per runa loop or retry storm burning moneya call and token cap that stops the run and hands it to a person
Checkpoints and resumea crash losing work, or redoing paid callsstate written after every node; load() continues from node
Idempotent side effectsduplicate emails or refunds after a resumean outbox keyed by a stable hash of ticket and reply
Human interruptmoney moving without approvalawaiting_approval state, resumed after a person decides
Tracenot knowing where or why runs failone JSON line per node with the trace id

Discussion

  1. Framework or hand-written? Our runner is about thirty lines; LangGraph, Temporal and the vendor SDKs add persistence, retries, visualisation and hosting. A framework pays for itself when you need durable waits of hours or days and many workflows; the concepts are the same either way.
  2. Where should budgets live? In code, per run, enforced before each call, as in the program. A budget written only in a prompt is a suggestion.
  3. What is the first thing to break when the workflow grows? Usually the interfaces: a schema changes in one step and the next step still expects the old one. Version schemas and put the version in the trace.

4.8 Workflows in production: what companies built and what they report

This section reads engineering write-ups for the same four things each time: which pattern, why, what went wrong, and what they report. Every quotation was checked against the live page. Each case separates what the company reports from our reading of it, and attributes numbers to the exact system they describe. Where a page gives no number, none is given here. Pages that block automated browsers are quoted with a link rather than screenshotted.

Workflow patterns in the write-ups of Section 4.8 (filled = stated on the page)chainroutingsectioningvoting/multi-runorchestratorevaluatorhuman gateUber uReviewUber uReview: chainUber uReview: routingUber uReview: sectioningUber uReview: voting/multi-runUber uReview: orchestratorUber uReview: evaluatorUber uReview: human gateUber QueryGPTUber QueryGPT: chainUber QueryGPT: routingUber QueryGPT: sectioningUber QueryGPT: voting/multi-runUber QueryGPT: orchestratorUber QueryGPT: evaluatorUber QueryGPT: human gateUber Genie (EAg-RAG)Uber Genie (EAg-RAG): chainUber Genie (EAg-RAG): routingUber Genie (EAg-RAG): sectioningUber Genie (EAg-RAG): voting/multi-runUber Genie (EAg-RAG): orchestratorUber Genie (EAg-RAG): evaluatorUber Genie (EAg-RAG): human gateLinkedIn assistantLinkedIn assistant: chainLinkedIn assistant: routingLinkedIn assistant: sectioningLinkedIn assistant: voting/multi-runLinkedIn assistant: orchestratorLinkedIn assistant: evaluatorLinkedIn assistant: human gateLinkedIn SQL BotLinkedIn SQL Bot: chainLinkedIn SQL Bot: routingLinkedIn SQL Bot: sectioningLinkedIn SQL Bot: voting/multi-runLinkedIn SQL Bot: orchestratorLinkedIn SQL Bot: evaluatorLinkedIn SQL Bot: human gateDoorDash Dasher supportDoorDash Dasher support: chainDoorDash Dasher support: routingDoorDash Dasher support: sectioningDoorDash Dasher support: voting/multi-runDoorDash Dasher support: orchestratorDoorDash Dasher support: evaluatorDoorDash Dasher support: human gateSpotify HonkSpotify Honk: chainSpotify Honk: routingSpotify Honk: sectioningSpotify Honk: voting/multi-runSpotify Honk: orchestratorSpotify Honk: evaluatorSpotify Honk: human gateAirbnb test migrationAirbnb test migration: chainAirbnb test migration: routingAirbnb test migration: sectioningAirbnb test migration: voting/multi-runAirbnb test migration: orchestratorAirbnb test migration: evaluatorAirbnb test migration: human gateStripe MinionsStripe Minions: chainStripe Minions: routingStripe Minions: sectioningStripe Minions: voting/multi-runStripe Minions: orchestratorStripe Minions: evaluatorStripe Minions: human gateAnthropic ResearchAnthropic Research: chainAnthropic Research: routingAnthropic Research: sectioningAnthropic Research: voting/multi-runAnthropic Research: orchestratorAnthropic Research: evaluatorAnthropic Research: human gateMicrosoft Magentic-OneMicrosoft Magentic-One: chainMicrosoft Magentic-One: routingMicrosoft Magentic-One: sectioningMicrosoft Magentic-One: voting/multi-runMicrosoft Magentic-One: orchestratorMicrosoft Magentic-One: evaluatorMicrosoft Magentic-One: human gateChains dominate; most add routing in front or an evaluator behind; orchestrators appear where thedecomposition cannot be written in advance (open-ended research). Voting/multi-run at Uber is re-running review five times to evaluate.
Workflow patterns stated in the production write-ups of this section, one row per system. Chains are the most common shape, usually with a router in front or an evaluator behind; orchestrator-workers appear for open-ended research. Filled cells are patterns the page describes, not our inference.

Uber: a chain of filters, and a router of intents

Anthropic: orchestrator-workers for research

Section 4.5 read the post in detail. What Anthropic reports: a lead agent with parallel subagents "outperformed single-agent Claude Opus 4 by 90.2%" on an internal research eval; multi-agent systems "use about 15× more tokens than chats"; early versions spawned "50 subagents for simple queries"; parallelism "cut research time by up to 90% for complex queries"; the lead agent waits for each set of subagents, which "creates bottlenecks". For evaluation they use a single model judge with a rubric (factual accuracy, citation accuracy, completeness, source quality, tool efficiency) scoring 0.0 to 1.0 with a pass-fail grade, plus human testing, starting from "a set of about 20 queries". Our reading: the pattern paid where subtasks were independent and answers valuable; the failures were orchestration failures (effort, briefs, stopping) fixed in the orchestrator's prompt and in code limits.

DoorDash: a guarded pipeline for support

DoorDash's September 2024 post "Path to high-quality LLM-based Dasher support automation" (DoorDash's site blocks automated browsers, so it is quoted here) describes a support system for delivery drivers. What DoorDash reports: a fixed retrieval pipeline with a two-tier guardrail (a cheap semantic-similarity check, then a model-based evaluator) that can retry or hand the conversation to a person, plus an offline model judge on five quality dimensions. "This guardrail system has successfully reduced overall hallucinations by 90% and cut down potentially severe compliance issues by 99%"; the guardrail's latency "is a notable drawback", and a more sophisticated guardrail model was dropped because "increased response times and heavy usage of model tokens made it prohibitively expensive". A November 2025 post from the same company ("Beyond Single Agents") argues for starting with "deterministic workflows" and warns that "you can't jump straight to sophisticated, multi-agent collaboration". Our reading: a chain with an evaluator at the end, ordered cheap check first and model check second, with a human fallback rather than a long retry loop; the cost of the evaluator shaped the design.

More production cases: Stripe, Airbnb, LinkedIn, Uber Genie, Spotify, Microsoft and Amazon (optional reading)

Stripe and Airbnb: workflows written as state machines. Two coding pipelines describe themselves in this chapter's vocabulary. Stripe (February 2026, "Minions") reports that its background coding agents follow "blueprints", workflows "defined in code" that look like "a state machine that intermixes deterministic code nodes and free-flowing agent nodes", with at most two rounds of continuous integration "since CI runs cost tokens, compute, and time". Airbnb (March 2025, "Accelerating Large-Scale Test Migration with LLMs", quoted because the page blocks automated browsers) migrated about 3,500 test files with a per-file pipeline "modeled ... like a state machine", retrying steps with the validation errors "until they passed or we reached a limit"; it reports 75% of files migrated "in just four hours", 97% after four days of tuning, and the remaining 3% finished by hand. Our reading: both are gated chains with code evaluators (tests, linters, CI) and capped retries, with a model inside some nodes and people at the end of the tail.

Spotify, Honk (December 2025). Spotify's background coding agent ends with deterministic verifiers (build, tests, formatting) and then a model judge that compares the diff with the original prompt: "The judge is simple. It uses the diff of the proposed change and the original prompt, and sends them to an LLM for evaluation." The post reports that "the judge vetoes about a quarter" of thousands of sessions and that "the agent is able to course correct half the time", and states plainly: "We have yet to invest in evals for our judge." Our reading: an evaluator placed after the code checks, in the order this chapter recommends, whose own rr and ff are not yet measured.

LinkedIn, SQL Bot (December 2024) checks generated queries with validators that "access new information not available to the query writer" (tables and fields exist; EXPLAIN runs) and feeds errors to a correction step (Chapter 3 reads it in detail). Microsoft's Magentic-One (Section 4.5) is the research orchestrator with explicit task and progress ledgers. Amazon's code-transformation agent in Q Developer has developers "review and iterate on the plan before the agent implements it", then builds and tests the result (Chapter 3). We could not verify public engineering write-ups with workflow details for Netflix, Instacart or Klarna, so they are not included.

The table

Company, systemPatternWhy (as stated)Reported problem or limitReported result
Uber, uReviewsectioned generation, then a grader and filterssingle prompts gave false positives and low-value commentsnoise; thresholds need per-language tuning75% of comments marked useful; 65% addressed
Uber, QueryGPTintent router, then a chain; human edits the tablesaccuracy fell as tables grewhallucinated tables and columns; ~5% run-to-run variancenot given as a single number
Uber, Genie (EAg-RAG)chain of pre- and post-processing agents; offline judgeincomplete or wrong answers; slow expert evaluationevaluation took experts weeks+27% relative acceptable answers, -60% relative incorrect advice
LinkedIn, Premium assistantrouter, retrieval, generationdifferent difficulty per step~10% of structured outputs invalid; last 15% of quality slow80% of target in one month; four more months towards 95%
DoorDash, Dasher supportpipeline with a two-tier guardrail and human fallbackhallucination and compliance riskguardrail latency; a stronger guardrail too costly-90% hallucinations, -99% severe compliance issues
Spotify, Honkagent with code verifiers, then a model judgeagents straying outside the promptjudge not yet evaluatedjudge vetoes about a quarter; half recover
Stripe, Minionsstate machine of code nodes and agent nodes; CI as evaluatorCI rounds cost tokens and timediminishing returns from more CI roundscapped at two CI rounds
Airbnb, test migrationper-file state machine with capped retriesa known transformation over 3,500 filesa long tail automation could not fix75% in four hours; 97% in four days; rest by hand
Anthropic, Researchorchestrator-workers with parallel subagentsbreadth-first questions15x chat tokens; over-spawning; synchronous bottleneck+90.2% on an internal eval

Three things hold across the table, as patterns rather than laws. Most systems are chains, with a router in front or an evaluator behind; the orchestrator appears where the question is open-ended and valuable. Evaluators are usually placed after cheap code checks, and their cost and latency shaped the designs (DoorDash dropped a stronger guardrail; Stripe capped CI rounds). And the teams that report the most progress also report how they measured each step, not only the end result.

Discussion

  1. Why are most production systems chains? The tasks automated first are the ones whose steps are known, and a known chain is cheaper to run, test and explain. That is a reasonable order of work, not a lack of ambition.
  2. Which reported numbers can you compare? Few. Each company measures its own system on its own data with its own definition of success (useful comments, acceptable answers, vetoed sessions). Read them as evidence that the pattern was worth keeping, not as benchmarks.
  3. What is missing from the write-ups? Mostly the evaluators' own error rates: how often the grader, guardrail or judge is wrong. Spotify says so openly; the rest are silent. That is the measurement to add first in your own system.

4.9 Choosing and evolving a workflow

The decision is not a single choice but an order of additions, each justified by a measured problem. The questions below map task properties to patterns.

Task propertyHow to measure itWhat it argues for
Variability of inputscluster a sample of real inputs; count the kindsone kind: a chain; several kinds with different handling: a router in front
Independence of subtasksfor each pair of subtasks, does one need the other's output?independent: parallel sections; dependent: a chain; unknowable in advance: an orchestrator
Verifiabilitylist the checks you can write in code, and the ones that need judgementcode checks: gates and an evaluator loop; judgement only: a measured judge, or a person
Latency budgetthe product's p95 target against the chain's sequential depthtight: parallelise, use smaller models per step, cut steps
Cost per tasktokens per run from traces, times traffichigh volume: routers and cascades; high value per task: orchestrators may pay
Blast radiuswhat happens if a wrong output ships: a reply, a refund, a deleted recordhigh: a human interrupt before the side effect, and idempotent steps
Evolving a workflow: start with a chain, add one pattern per measured problemstartyesa single call, then a chain of fixed stepsgates between steps; an eval set for each stepnoinputs diverge?yesadd a router in frontmeasure its confusion matrix; keep an unknown bucketnolatency too high?yesrun independent sections in parallelcost stays the sum; time becomes the slowestnosplit not knowable?yesadd an orchestrator over workerscost: every worker's loop; write precise briefsnoquality measurable?yesadd an evaluator loop with a stop rulecode checks first; measure the judge firstnosteps unknowable?yesonly then an agent loop (Chapter 3)a budget, a trace and a verifier still applyre-measure accuracy, cost per correct answer and p95 latency after every addition
A migration path, as a starting heuristic rather than a fixed sequence. Start with a gated chain; add a router when inputs diverge, parallel sections when latency hurts, an orchestrator when the split cannot be written in advance, an evaluator where quality can be measured, and an agent loop only when the steps themselves cannot be known. Re-measure after every addition.

The migration path in the figure is a heuristic: the patterns compose, so a system may need an evaluator before it needs a router, and some tasks start as orchestrators. What does hold is the order of evidence. Each addition should answer a failure you have seen in traces and labelled data, and each should be re-measured on accuracy, cost per correct answer and latency once it is in.

The one-rule summary of the chapter: hand the model a decision only when you cannot write it in code, and measure every decision you hand over. A chain hands over none; a router hands over one, measurable with a confusion matrix; parallel calls hand over none and buy time; an orchestrator hands over the shape of the work, measurable by grading plans and counting duplicated effort; an evaluator hands over the decision to stop, measurable by its false-accept and false-reject rates. The experiments in this chapter found the same thing from four directions: the model was good at the content of each step and unreliable at the decisions around it (whether its answer was good enough, whether a draft had 55 words, whether a question was hard), and code made those decisions better wherever code could make them.

Discussion

  1. Where should Northwind start? With the gated chain of Section 4.2, a labelled set of a few hundred tickets, and per-step evals. Everything else waits for a measured reason.
  2. When is it right to skip a step on the path? When the task's structure is obvious: code generation against tests goes straight to an evaluator loop with the tests as the evaluator; a research question goes straight to an orchestrator. Skipping is fine if you still take the measurement that would have justified the step.
  3. What should be re-measured when the model is upgraded? Every model decision: the router's confusion matrix, the judge's error rates, the per-step success rates. A new model can improve the content and change the decisions in either direction.

Exercises

  1. Northwind's chain has three steps at 95%, 92% and 98% per step on a labelled set. Compute the end-to-end success under the independence assumption, say which step a gate would help most, and describe one way the independence assumption could fail on real tickets.
  2. A cascade's check accepts 90% of right cheap answers and 25% of wrong ones. Compute the precision of accepted answers when the cheap model is right 80% of the time and when it is right 40% of the time, and decide in which case you would ship the cascade.
  3. Design the brief an orchestrator should give one worker for "why did delivery complaints rise last quarter?": objective, output format, tools and sources, boundaries. Then list two measurements that would tell you whether workers are duplicating each other.
  4. Your evaluator loop uses a model judge for tone and a word count in code. Traces show the loop runs four rounds on 30% of drafts. Describe how you would find out whether the judge is producing false fails, and what you would change if it is.
  5. In ch4_workflow.py, the send step is idempotent because of its outbox key. Name two other side effects a support workflow might have, and how you would make each idempotent.

Key takeaways

  • A workflow is model calls arranged by control flow the engineer wrote; an agent lets the model choose the next step. Most production systems are workflows with a model inside some nodes.
  • Draw the system as a graph of nodes and edges, and mark which edges code decides and which the model decides. Each decision handed to the model needs its own measurement.
  • Prompt chains make each step easy and testable. End-to-end success is roughly the product of per-step success, so gates between steps matter: in our run, gates with one retry lifted a loosely specified chain from 12 to 16 of 20, and could not catch valid-but-wrong values.
  • Routers are judged by their confusion matrix, not their accuracy. Our model router sent one of 24 questions to the strong model and caught none of the seven that needed it; a cascade reached 88% at about two-thirds of the strong model's cost per correct answer.
  • The false-accept rate of a check is not one minus its precision: precision depends on the base rate, p r/(p r+(1−p) f)p\,r / (p\,r + (1-p)\,f).
  • Parallel calls buy time, not tokens (4.5 times faster in our run). A plurality vote needs the right answer to be the most frequent, not a majority; correlated errors cut the gain, and our vote gained nothing. Best-of-n with a verifier helped, by less than independence predicts.
  • Orchestrator-workers fits open-ended, breadth-first, valuable tasks. Anthropic reports about 15 times the tokens of a chat; the failures they report are about briefs, effort and stopping.
  • In an evaluator-optimiser loop, put code checks first and measure any judge before it stops a loop. Our judge agreed with a word count 32% of the time, and its false fails both cost rounds and damaged a good draft.
  • Production plumbing is part of the pattern: typed state, schema checks, budgets, checkpoints, idempotent side effects, human interrupts and one trace line per node, all shown working in ch4_workflow.py.
  • Hand the model a decision only when you cannot write it in code, and re-measure every handed-over decision after each change.

References

Papers

  1. Tongshuang Wu, Michael Terry, Carrie J. Cai. AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. CHI 2022 (arXiv October 2021).
  2. Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, Ed Chi. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. ICLR 2023 (arXiv May 2022).
  3. Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, Ashish Sabharwal. Decomposed Prompting: A Modular Approach for Solving Complex Tasks. ICLR 2023 (arXiv October 2022).
  4. Lingjiao Chen, Matei Zaharia, James Zou. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv May 2023.
  5. Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, Ahmed Hassan Awadallah. Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing. ICLR 2024 (arXiv April 2024).
  6. Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, Ion Stoica. RouteLLM: Learning to Route LLMs with Preference Data. ICLR 2025 (arXiv June 2024).
  7. Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, Deheng Ye. More Agents Is All You Need. TMLR 2024 (arXiv February 2024).
  8. Xinyun Chen, Renat Aksitov, Uri Alon, Jie Ren, Kefan Xiao, Pengcheng Yin, Sushant Prakash, Charles Sutton, Xuezhi Wang, Denny Zhou. Universal Self-Consistency for Large Language Model Generation. arXiv November 2023.
  9. Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, Azalia Mirhoseini. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. arXiv July 2024.
  10. Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, Yueting Zhuang. HuggingGPT (full title on arXiv). NeurIPS 2023 (arXiv March 2023).
  11. Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, Chi Wang. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv August 2023.
  12. Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Erkang Zhu, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, Peter Chang, Ricky Loynd, Robert West, Victor Dibia, Ahmed Awadallah, Ece Kamar, Rafah Hosn, Saleema Amershi. Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks. arXiv November 2024.
  13. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, Ion Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks (arXiv June 2023).
  14. Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, Chenguang Zhu. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. EMNLP 2023 (arXiv March 2023).
  15. Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, Denny Zhou. Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024 (arXiv October 2023).

Engineering blogs and docs

  1. Anthropic (Erik Schluntz, Barry Zhang). Building effective agents. 19 December 2024.
  2. Anthropic (Jeremy Hadfield, Barry Zhang, Kenneth Lien, Florian Scholz, Jeremy Fox, Daniel Ford). How we built our multi-agent research system. Engineering blog, 13 June 2025.
  3. OpenAI. A practical guide to building agents. PDF, April 2025.
  4. LangChain. Graph API overview and Interrupts. LangGraph documentation, accessed October 2026.
  5. Temporal. Understanding Temporal. Temporal documentation, accessed October 2026.
  6. AWS (Aaron Sempf, Andrew Hooker). Agentic AI patterns and workflows on AWS: Workflow for routing. AWS Prescriptive Guidance, July 2025.
  7. Uber. uReview: Scalable, Trustworthy GenAI for Code Review at Uber. Engineering blog, 12 August 2025.
  8. Uber. QueryGPT: Natural Language to SQL Using Generative AI. Engineering blog, 19 September 2024.
  9. Uber. Enhanced Agentic-RAG: What If Chatbots Could Deliver Near-Human Precision?. Engineering blog, 29 May 2025.
  10. LinkedIn (Juan Pablo Bottaro, Karthik Ramgopal). Musings on building a Generative AI product. Engineering blog, 25 April 2024.
  11. LinkedIn (Albert Chen and colleagues). Practical text-to-SQL for data analytics. Engineering blog, 9 December 2024.
  12. DoorDash. Path to high-quality LLM-based Dasher support automation. Engineering blog, 17 September 2024; and Beyond Single Agents: How DoorDash is building a collaborative AI ecosystem. 11 November 2025.
  13. Spotify (Max Charas, Marc Bruggmann). Background Coding Agents: Predictable Results Through Strong Feedback Loops (Honk, Part 3). Engineering blog, 9 December 2025.
  14. Stripe (Alistair Gray). Minions: Stripe's one-shot, end-to-end coding agents. stripe.dev blog, 9 February 2026.
  15. Airbnb (Charles Covey-Brandt). Accelerating Large-Scale Test Migration with LLMs. Airbnb Tech Blog, March 2025.
  16. AWS (Aytul Arisoy Cholkar). Amazon Q Developer just reached a $260 million dollar milestone. AWS DevOps blog, 1 August 2024.

Code for this chapter

  1. code/agents/ch4_chain.py: the three-step ticket chain under four conditions (strict or loose first step, with or without gates) against gpt-5-mini; results in results/ch4_chain.json.
  2. code/agents/ch4_router.py, ch4_router_stats.py and ch4_grade.py: always-cheap, always-strong, cascade and classifier router on 24 questions, the cascade check read as a verifier, and the typed grader with its adversarial self-test.
  3. code/agents/ch4_parallel.py: plurality voting, best-of-n with a programmatic verifier, and sequential against concurrent wall-clock time.
  4. code/agents/ch4_evaluator.py: the evaluator-optimiser loop with a code checker or a model judge, judge-checker agreement, and a held-out fact check.
  5. code/agents/ch4_cost.py: chain reliability, routing cost, plurality voting with correlated errors, orchestrator cost, and the precision of a stop with an imperfect evaluator, each with its assumptions printed.
  6. code/agents/ch4_workflow.py: the complete composed workflow with typed state, schema-validated steps, budgets, checkpoints, resume, a human interrupt, an idempotent send and a trace; ch4_llm.py is the shared API helper.
  7. code/agents/figs_ch4.py and code/agents/shots_ch4.py: the figures (with an automatic text-overlap check) and the paper excerpts.

Next

→ Chapter 5: Multi-agent systems

This chapter kept the control flow in code wherever it could and handed the model one decision at a time. Chapter 5 follows the orchestrator-workers pattern to its end: systems of several agents, each with its own loop, that hand work to one another, debate or vote. It asks the questions this chapter asked of every pattern (what changes, what it costs, how it fails and how you test it) of systems where the graph itself is partly written by the models, and it reads the evidence on when several agents beat one and when they only multiply the bill.