Agentic Systems · Part 2 · Patterns
Chapter 4 · Workflow patterns: chaining, routing, parallel work, orchestrators and evaluators
Workflow patterns: prompt chains, routers, parallel calls, orchestrators and evaluators; what each costs, how each fails and how to test it.
Goal: by the end of this chapter you can draw any LLM system as a graph of steps, say which steps your code decides and which the model decides, and choose among the five workflow patterns (prompt chaining, routing, parallelisation, orchestrator-workers, evaluator-optimiser) by what each costs, how it fails and how you will test it. One running example carries you through all five, four small experiments against a real model show where each pattern earns its keep, and one complete program puts them together with the plumbing production needs: typed state, schema checks, budgets, retries, checkpoints, a human approval step and a trace.
4.1 Workflows versus agents: who holds the control flow
The running example: Northwind Home's support inbox
Northwind Home is a fictional online shop that sells furniture and kitchenware in the UK. Its support inbox receives a few thousand messages a week: "order #48213 arrived with the screen cracked", "I was charged twice for ORD-10988", "do you ship to the Isle of Man?", "the sofa from order 61234 has a torn seam, I want my money back, £649". The team wants a model to read each message, work out what it is about, draft a reply, check the reply against company policy, and send it, with a person approving anything that costs money.
The first version is one prompt: "here is the policy, here is the message, write the reply". It mostly works, and it fails in ways nobody can see from outside: the order number is quoted wrongly, the policy's ban on promising delivery dates is ignored, a refund is offered on a question about shipping zones. Every fix is another sentence in a prompt that keeps growing, and every new sentence can break a case that used to work.
This chapter rebuilds that system five times, once per pattern, and asks the same four questions each time: what changes, what it costs, how it fails, and how you test it. Underneath all five is one question: who decides what happens next?
Anthropic's December 2024 post "Building effective agents" drew this line, and the industry adopted it. It names five workflow patterns, and Sections 4.2 to 4.6 take them in turn.
Read the figure as a sequence of hand-overs. A chain hands the model no decisions about flow. A router hands it one, made once: which branch should this input take? Parallel sections hand it none; code fans the work out and merges it. An orchestrator hands it the shape of the work: how many subtasks, and which. An evaluator loop hands it the decision to stop. An agent hands it every step. Each hand-over buys adaptability and usually costs predictability, testability, money and time. So a useful default is to hand the model a decision only when you cannot write that decision in code, and to know which decisions you have handed over.
Graphs: the vocabulary of workflows
Every workflow can be drawn as a graph, and the frameworks that run them (LangGraph, the OpenAI Agents SDK, Temporal, AWS Step Functions) all use graph words.
Research notes: OpenAI's view from the agent end (optional reading)
OpenAI's practical guide of April 2025 frames the same choice from the other side: start with one agent and add structure only when it fails.
The two vendors start at opposite ends and meet in the middle: add a moving part when a measurement asks for it.
Discussion
- Is a router an agent? The model makes a decision about flow. My view: no, as long as the branches are fixed in code and the router can be tested against labels like any classifier; it becomes agent-like when the model can name a branch you did not write.
- Does the distinction matter to users? They see an answer, not a graph. But latency ceilings, cost ceilings and testability are properties of the graph, and those are what break in production.
- Will better models make workflows obsolete? Better models move the line towards agents for open-ended tasks. For high-volume, well-understood tasks, a fixed path's predictability and cost per call are part of the product, so the line moves more slowly there.
4.2 Prompt chaining: fixed steps, with gates in between
The example: split Northwind's one prompt into three
The one-prompt system drops requirements because it is asked to do four things at once: understand the ticket, choose a response, write it, and obey the policy. A prompt chain gives each job its own call:
- Extract: read the ticket and return JSON with the issue type (one of five), the order id in the form
ORD-NNNNN, and the urgency. - Draft: from those facts, choose the reply category from a fixed table and write the reply.
- Check: test the reply against the policy (at most 70 words, quote the order id, no promises of timing such as "as soon as possible", no talk of refunds outside billing).
Between the steps sit gates: code that checks the previous step's output before the next one runs.
Why chains need gates: the arithmetic
A chain multiplies probabilities. If each step returns a usable, correct output with probability , the whole chain of steps succeeds with probability
Assumptions: steps fail independently, a bad output anywhere ruins the result, and no later step repairs an earlier mistake. Real chains break all three in both directions: one confusing ticket makes several steps fail together, and a good drafting step can paper over a slightly wrong extraction. Treat the formula as a way to see the shape, not a forecast. The shape is unforgiving: five steps at 90% succeed 59% of the time; ten succeed 35%.
Now add a gate that detects a fraction of bad outputs and retries the step once. The step succeeds if the first attempt is good, or if it is bad, the gate notices and the retry is good:
Assumptions: the retry succeeds with the same probability as the first attempt (in practice the error message often helps, so this is a floor), and the gate never rejects a good output. With and , , and five steps succeed 87% of the time instead of 59%, for about 8% more calls. The large condition is : the share of bad outputs the gate can see. A gate sees a missing key, an unknown label, a malformed id, a reply over the word limit. It cannot see a well-formed answer that is wrong.
Our experiment: Northwind's chain, with and without gates
The script ch4_chain.py runs this chain on twenty short tickets with a real model (gpt-5-mini, reasoning effort set to minimal). The tickets are deliberately messy: order numbers written five ways ("order #48213", "ORD 55501", "order no. 30377"), one in Spanish, one mentioning two orders, several with no order at all. Each ticket has gold labels: its issue type and its canonical order id.
There are four conditions, varying one factor at a time. The step-1 prompt is either strict (it names the keys, the five allowed values and the ORD-NNNNN format) or loose ("extract the issue type, the order id and the urgency ... reply in JSON", the kind of prompt a first version often has). Each runs without gates or with gates (a schema check after step 1, a schema and policy check after step 2, one retry each). Success has two parts, reported separately because they differ in independence:
- Gold correct: the right reply category and the gold order id quoted in the reply. No gate ever sees the gold labels, so this is an independent measure.
- Policy correct: word limit, no timing promises, no refund talk outside billing. Gate 2 checks the same rules, so in the gated conditions this part is graded by the same code that gave the feedback. It shows compliance, not independent quality.
The whole run cost about four US cents; token counts come from the API's usage fields.
# code/agents/ch4_chain.py (abridged): one chain step with a gate. Without gates, whatever came back flows on.
def step(prompt, gate, gated, use):
text, u = llm.ask(prompt, max_tokens=500) # one model call; usage recorded from the API
use.append(u)
d = llm.parse_json(text)
if not gated:
return d, text, 0
problems = gate(d) # e.g. "order_id must match ORD-NNNNN or be null, got '48213'"
if not problems:
return d, text, 0
retry = prompt + '\n\nYour previous answer failed these checks:\n- ' + '\n- '.join(problems) + \
'\nReply again with a corrected JSON object only.'
text, u = llm.ask(retry, max_tokens=500) # one retry, with the exact failures in the prompt
use.append(u)
d = llm.parse_json(text)
return (d if not gate(d) else None), text, 1 # None stops the chain: hand the ticket to a personWhen the prompt is specific, the gates are idle. With the strict prompt the model returned valid labels and canonical ids on every ticket; the gates never fired, and both conditions scored 19 of 20. The one failure was "Do you ship to the Isle of Man?", labelled "delivery" instead of "other". That is a valid label, so no gate could catch it; it needs a labelled set to find.
When the prompt is loose, the gates do the work. With the loose prompt every step-1 output parsed as JSON and almost none was usable: issue types came back as "Damaged item / cracked screen" and "Incorrect VAT rate on invoice", order ids as "48213" and "ORD 55501". Without gates these flowed on and 8 of 20 replies went to the wrong category or failed to quote the order id. With gates, all twenty tickets failed their first check, were retried once with the exact problems listed, and 16 of 20 succeeded. Of the four failures, one was the Isle of Man label, one was a ticket the gate correctly stopped and handed to a person (it mentioned two orders), and two are the instructive ones: the retry "fixed" the order id by setting it to null, which the schema allows for tickets with no order. A retry optimises for the gate, so a gate with an easy wrong exit gets that exit.
A model is a poor checker of mechanical rules. Step 3, the model asked whether the reply broke the policy, disagreed with a plain-code check on most tickets. Across the four conditions it flagged "over 70 words" 71 times on replies that a word count passed. A rule you can write in code (a word count, a list of banned phrases, a regular expression) is better checked in code, and a model checker belongs where code cannot reach, such as tone, after you have measured it (Section 4.6).
The research behind chaining
The idea was first studied as a question of human control. In October 2021 Tongshuang Wu, Michael Terry and Carrie Cai at Google posted "AI Chains" (CHI 2022). Large models of the time did well on single operations and poorly on tasks with several parts, and users could not see or fix what went wrong inside one big prompt.
The study's measures were participants' ratings, not accuracy; the authors note the limits (twenty people, two tasks, a 2021 model). In 2022 two papers turned decomposition into a reasoning method: Least-to-Most prompting (Zhou et al., Google) and Decomposed Prompting (Khot et al., Allen Institute for AI). Both found the largest gains on problems with many steps, and both point at the same practical conclusion: the decomposition is the hard part, and when an engineer can write it, it can live in code.
Research notes: Least-to-Most, Decomposed Prompting and the AI Chains abstract (optional reading)
What a production team learned about gates
Prompt chaining in practice
| One big prompt | Prompt chain with gates | |
|---|---|---|
| Pros | one call, lowest latency; nothing to wire | each step easy for the model and testable alone; different models per step; format failures caught in code; one step changes without touching the others |
| Cons | requirements dropped silently; nothing to test between input and output | latency adds up; errors propagate; interfaces between steps must be designed and versioned; a gate can push a retry into a wrong-but-valid answer |
| When to pick it | the task is one operation (classify, extract, rewrite) | the task has distinct operations in a known order, especially if some can be code |
| Production failure | a missed requirement nobody notices until a user does | an upstream change (new input formats) degrades every step downstream |
What to measure before you decide
| Measurement | How | What it tells you |
|---|---|---|
| Per-step success rate | label each step's output on a hundred inputs | the chain's ceiling; which step to improve first |
| Gate detection rate | label the failures the gate missed | how much a gate can recover; valid-but-wrong is the rest |
| First fault position | for each failed run, the first step whose output was wrong | whether to fix extraction or drafting |
| Retry outcome | for each retry: fixed, still wrong, wrong in a new valid way | whether the gate's message helps or opens an easy exit (our null ids) |
| Latency per step | wall clock per node from traces | which step to shrink, parallelise or give a smaller model |
Discussion
- How many steps should a chain have? One per distinct operation, and no more. Merge two model steps whose combined prompt is still a single operation, and replace any step that can be code.
- Should a step see the original input or only the previous output? Passing only structured outputs keeps contexts small but loses information. A reasonable rule: steps that write for a person see the original input; steps that classify or check see only what they need.
- Retry or repair? LinkedIn chose a parser for latency. My view: repair what you can predict, retry what you cannot, and count both in the trace, because a rising repair rate is often the first sign that an upstream prompt or model has changed.
4.3 Routing: classify, then dispatch
The example: not every Northwind question needs the expensive model
Northwind's team also runs an internal help desk where staff ask the model quick questions. Most are easy ("what is the plural of criterion?", "what is 15% of 240?"); some are multi-step calculations of the kind that price disputes produce ("3 pens for £2.40 and notebooks at £1.75; what do 9 pens and 4 notebooks cost?"). Sending everything to the strongest model is accurate and expensive; sending everything to the cheapest is cheap and wrong on the hard ones. A router looks at each input and sends it to the handler that suits it.
The arithmetic: a router's confusion matrix is its cost model
A router makes two kinds of mistake. It sends a hard input to the cheap handler (accuracy is lost), or an easy input to the expensive one (money is lost). Write for the share of hard inputs, for the router's recall on hard inputs (the share of hard inputs it escalates) and for its false-escalation rate (the share of easy inputs it escalates). Then the share of traffic that reaches the strong model is
and the cost per question is the router's own cost plus the cheap and strong costs weighted by their shares. Assumptions: the two error rates do not depend on the input, and each model's accuracy on easy and hard inputs is fixed. ch4_cost.py tabulates this with illustrative accuracies; the point it makes is that every point of lost recall costs accuracy and every point of false escalation costs money, so you cannot judge a router by "accuracy" alone. You need its confusion matrix.
Our experiment: a cascade and a classifier router
ch4_router.py answers 24 questions with exact answers (a mixed set standing in for the help desk: eight lookups, eight multi-step calculations, eight trickier counting and date problems) in five ways: always-cheap, always-strong, a cascade (two cheap samples; if they agree, accept, otherwise ask the strong model), a classifier router (one cheap call labels the question easy or hard), and an oracle router that knows which questions the cheap model gets right. "Cheap" is gpt-5-mini at minimal reasoning effort; "strong" is the same model at medium effort, so the price per token is identical and the difference is the hidden reasoning tokens the medium setting spends. That keeps the comparison on one verified price, and it means our quality gap is smaller than two different models would show. The five conditions differ only in the routing policy.
Grading is typed and exact. An earlier draft of this book graded by substring, which gives "184" credit for 84; ch4_grade.py instead requires exactly one number and compares it numerically, compares times as minutes, and compares words after explicit normalisation. It prints its own test on adversarial strings:
# code/agents/ch4_router.py (abridged): the two routing policies, computed from the same samples so they differ only in the policy.
def cascade(r): # two cheap samples; accept if they agree, otherwise pay for the strong model
if r['agree']:
return r['cheap_ok'], [r['u_cheap'], r['u_cheap2']]
return r['strong_ok'], [r['u_cheap'], r['u_cheap2'], r['u_strong']]
def router(r): # one cheap classification call decides before anything is answered
if r['route'] == 'easy':
return r['cheap_ok'], [r['u_router'], r['u_cheap']]
return r['strong_ok'], [r['u_router'], r['u_strong']]Three results, each a general point.
The classifier router failed in the way routers usually fail: it was overconfident. Asked "would a small model handle this?", the cheap model said yes to 23 of 24 questions, including all seven it then got wrong. Its recall on the hard questions was zero, so it matched always-cheap's 71% while paying for an extra call on every question. A router's error is invisible in end-to-end accuracy until you build the confusion matrix, and here the matrix is the whole story.
The cascade worked, and its check was weaker than it looked. It escalated 5 of 24 questions and reached 88% (21 of 24) at $0.135 per thousand correct answers, against $0.194 for always-strong. But read its agreement check as a verifier, which ch4_router_stats.py does from the saved results. The check accepted 16 of the 17 right cheap answers ( = 94%) and also 3 of the 7 wrong ones ( = 43%): on "what is the sum of the digits of 2 to the power 20?", both cheap samples said 7 (the answer is 31). When a model is confidently wrong, it is wrong twice.
The false-accept rate is not one minus precision. Precision, the share of accepted answers that are right, depends on how often the cheap model is right in the first place. With the share of right cheap answers, Bayes' rule gives
For our cascade, , and give 84%: 16 of the 19 accepted answers were right. The same check on a task where the cheap model is right 30% of the time would give much lower precision. Assumptions: and are properties of the check that do not change with the mix of questions, which is only roughly true.
The research behind routing
Routing between models of different price became a research topic in 2023, when API prices differed by two orders of magnitude. Lingjiao Chen, Matei Zaharia and James Zou at Stanford posted FrugalGPT in May 2023 and made the cascade the reference design.
Two later papers trained routers that decide before any answer is generated. Microsoft's Hybrid LLM (Ding et al., ICLR 2024) routes between a small and a large model by predicted difficulty and reports "up to 40% fewer calls to the large model, with no drop in response quality". RouteLLM (Ong et al., UC Berkeley, Anyscale and Canva, June 2024) trains routers on human preference data from Chatbot Arena and reports cost reductions of "over 2 times" at matched quality. Both report their gains as curves of quality against the share of calls sent to the strong model, which is the honest way to show a router: a single accuracy number hides the trade.
Research notes: RouteLLM, Hybrid LLM and the FrugalGPT savings table (optional reading)
AWS's prescriptive guidance on agentic patterns (July 2025) lists the same use cases from the practitioner's side: "triaging requests across a variety of tasks", inputs that "must be preprocessed or normalized before entering more specialized workflows", and an agent "acting as a conversational switchboard".
Routing in practice
| Classifier router (decide first) | Cascade (answer cheaply, then check) | |
|---|---|---|
| Pros | one decision, then a specialised handler with its own prompt, model and eval set; can send "unknown" to a person | needs no labelled routing data to start; the strong model is paid for only on escalations |
| Cons | its mistakes are silent unless you build its confusion matrix; a model asked "is this hard?" tends to say no | every input pays for at least one cheap call; a weak check accepts confident wrong answers |
| When to pick it | inputs fall into distinct kinds with different prompts or tools | one kind of task, cheap and strong models of the same family, and a check you have measured |
| Production failure | a new kind of input is forced into the nearest branch | the cheap model drifts and the check keeps accepting |
What to measure before you decide
| Measurement | How | What it tells you |
|---|---|---|
| Router confusion matrix | label a few hundred inputs by outcome (did the cheap handler get it right?) | lost accuracy and wasted cost, separately |
| Check and | for a cascade, how often the check accepts right and wrong cheap answers | precision via Bayes; whether the cascade is safe |
| Strong-model share | the share of traffic escalated | cost per thousand requests, and how it moves as traffic changes |
| Unknown-bucket rate | inputs the router sends to a person or a fallback | coverage; a rising rate is new kinds of input arriving |
Discussion
- Rule, classifier or model? A rule is free and exact where a field decides; a trained classifier is cheap and measurable; a model handles the long tail. Most production routers are all three in that order.
- Can the cheap model be its own router? Ours said yes to 23 of 24 questions. My view: only if you have measured its self-assessment against outcomes; a separately trained scorer, as in FrugalGPT, is usually better.
- Where does a router go in a chain? Usually first, so each branch can be a simpler chain. A second router deep in a chain is often a sign the first one is too coarse.
4.4 Parallelisation: split the work, or do it several times
The example: two different reasons to make several calls at once
Before a refund is approved, Northwind wants three checks on the customer's message: does the refund fit the returns policy, is there a fraud signal, is the message abusive? They are independent, so there is no reason to run them one after another. That is sectioning. Separately, the shop's price calculations ("9 pens at 3 for £2.40 and 4 notebooks at £1.75, paid with a £20 note: what change?") are sometimes wrong; asking several times and taking the most common answer might help. That is voting.
The arithmetic: what voting can and cannot do
Suppose each sample is right with probability , and when it is wrong it picks one of different wrong answers equally often. A plurality vote returns the most frequent answer. It does not need a majority: a right answer given 40% of the time beats two wrong answers given 30% each. ch4_cost.py computes the exact probability that the plurality is right:
| right answer | distinct wrong answers | n = 1 | n = 5 | n = 11 |
|---|---|---|---|---|
| 0.4 | 1 (binary) | 40% | 32% | 25% |
| 0.4 | 2 | 40% | 45% | 50% |
| 0.4 | 5 | 40% | 58% | 76% |
| 0.6 | 2 | 60% | 77% | 90% |
Assumptions: samples are independent, is constant, and wrong answers spread evenly. Two consequences follow. Voting helps when the right answer is the most frequent one, and it helps more when wrong answers scatter; it hurts when one wrong answer is more common than the right one (the binary row). And real samples are not independent: if, with probability , all samples copy one shared misreading of the problem, accuracy becomes , which for , , falls from 45% at to 42% at . Correlated errors are why voting gains are usually smaller than the formula promises.
With a verifier that can check each candidate (a test, a calculation, a schema), you do not need the most frequent answer, only one that passes. If each sample passes with probability , at least one of passes with probability , under the same independence assumption and a verifier that never accepts a wrong answer.
Our experiment: voting, best-of-n with a verifier, and wall-clock time
ch4_parallel.py measures all three with the real model. Voting: ten multi-step word problems with one exact answer (Northwind's price-calculation kind), five samples each, plurality vote over the first 1, 3 and 5. Best-of-n with a verifier: ten "Game of 24" puzzles (combine four numbers with + − × ÷ into 24), chosen because a program can check any candidate exactly; a puzzle counts as solved if any of the first candidates verifies. Latency: five samples run one after another, then on five threads. The run cost under a cent.
# code/agents/ch4_parallel.py (abridged): the same five calls, one after another and then fanned out on threads.
t0 = time.time()
for _ in range(N):
llm.ask(P24_PROMPT.format(nums=nums), max_tokens=200) # sequential: latency is the sum
seq = time.time() - t0
t0 = time.time()
with ThreadPoolExecutor(max_workers=N) as ex: # fan-out: latency is the slowest call
list(ex.map(lambda _: llm.ask(P24_PROMPT.format(nums=nums), max_tokens=200), range(N)))
par = time.time() - t0 # the token bill is identical in both casesVoting did not help, and the samples show why. On the four problems the model solved, all five samples agreed (or four of five); on the six it missed, the samples scattered over wrong answers ("3.2", "6.15", "3.25", "5.2", "5.6" for a right answer of 5.80) or agreed on the same wrong one ("80" four times out of five for a right answer of 60). In neither case was the right answer the most frequent. Voting amplifies a model that is usually right on that question; it cannot create an answer the model rarely produces.
Best-of-n with a verifier helped, by less than independence predicts. Across all fifty candidates, 18% verified. If samples were independent, three would solve 45% of puzzles and five would solve 63%; we measured 30% and 30%. The verified candidates clustered on the same three puzzles, and on seven puzzles no candidate verified at all. Samples from one model at one temperature share their blind spots.
Fan-out bought time, not tokens. Five calls took 4.6 seconds one by one and 1.0 second on five threads, a 4.5 times speedup, for the same bill. For sectioning that is the whole point; for voting it means the latency of samples is about the latency of one.
The research behind parallel sampling
The research line runs from self-consistency (Chapter 3) through two 2024 papers that scaled sampling far further. Bradley Brown and colleagues at Stanford, Oxford and Google DeepMind asked in "Large Language Monkeys" (July 2024) what happens with hundreds or thousands of samples per problem.
Two other papers fill in the picture. "More Agents Is All You Need" (Li et al., Tencent, TMLR 2024) samples up to forty answers and votes, and finds gains that grow with task difficulty: with Llama2-13B, GSM8K accuracy went from 0.35 with one sample to 0.59 with forty, above Llama2-70B's single sample (0.54). Universal Self-Consistency (Chen et al., Google, November 2023) handles free-form answers, where exact voting is impossible, by asking a model to pick "the most consistent response based on majority consensus"; it matched exact-match voting on maths and execution-based voting on SQL without running the code.
Research notes: More Agents, Universal Self-Consistency and the Monkeys abstract (optional reading)
Parallelisation in practice
| Sectioning | Voting or best-of-n | |
|---|---|---|
| Pros | latency of the slowest section, not the sum; each section has a focused prompt and its own eval | can lift accuracy when samples vary and the right answer is common, or when a verifier exists |
| Cons | sections must really be independent, or the merge has to reconcile them; cost is the sum | times the tokens; correlated errors cut the gain; a weak aggregator picks wrong |
| When to pick it | separate checks or separate parts of a document | exact answers with high disagreement between samples, or any task with a cheap verifier |
| Production failure | two sections disagree and the merge silently picks one | a vote that is confidently wrong because every sample made the same mistake |
What to measure before you decide
| Measurement | How | What it tells you |
|---|---|---|
| Section independence | does any section's output change another's? | whether you can fan out at all |
| Agreement rate | sample the same input five times | if samples agree, one is enough; if the right answer is rare, voting cannot help |
| Per-sample pass rate with a verifier | verify every candidate on a labelled set | the ceiling for best-of-n, and how far correlation pulls you below |
| Wall-clock time, sequential against concurrent | time both on real traffic | the speedup your provider's rate limits actually allow |
Discussion
- Is sectioning just a chain with the arrows removed? Only when the sections do not depend on each other. If the fraud check needs the policy check's result, you have a chain, and running it in parallel creates a race.
- When is voting worth five times the tokens? When samples disagree often and the right answer is usually the most common one. Measure both on a hundred inputs first; our run would have saved 80% of its voting cost by checking agreement first.
- Different prompts or the same prompt? Varying the prompt or the model across samples reduces correlated errors, at the cost of comparability. Uber's code review (Section 4.8) uses different specialised assistants, which is sectioning, not voting.
4.5 Orchestrator-workers: when the split cannot be written in advance
The example: a question nobody can decompose in advance
Northwind's head of operations asks: "Delivery complaints rose by about forty per cent last quarter. Why, and what should we change?" No fixed chain answers this. The work depends on what the first look finds: if complaints cluster in two regions, someone has to read those regions' courier notices; if they cluster on one product, someone has to read its returns data. The number and kind of subtasks are decided by the input and by early findings. That is the case for an orchestrator: a model that plans the split, hands each part to a worker, and combines what comes back.
What the production evidence says
The clearest public account is Anthropic's description of the multi-agent system behind Claude's Research feature, published in June 2025.
The same post lists what went wrong early, in terms every orchestrator builder will recognise: agents "spawning 50 subagents for simple queries", "scouring the web endlessly for nonexistent sources", and subagents that "duplicated work" because the lead agent's briefs were short ("research the semiconductor shortage"). The fixes were in the orchestrator's prompt: briefs with "an objective, an output format, guidance on the tools and sources to use, and clear task boundaries", and explicit effort rules ("simple fact-finding requires just 1 agent with 3-10 tool calls ... complex research might use more than 10 subagents"). Running 3 to 5 subagents in parallel, each making several tool calls in parallel, "cut research time by up to 90% for complex queries".
The cost arithmetic, with its assumptions
ch4_cost.py compares one agent doing subtasks in sequence with an orchestrator that writes briefs, runs workers concurrently and synthesises. With a 1,500-token prompt, 300-token results and a 400-token brief, the orchestrator sends 20,600 input tokens at against 29,700 for the single agent, and finishes in 25 seconds against 72.
Assumptions, and why they matter: the single agent replays its whole growing history at every step (no compaction, no prompt caching, both of which cut its cost a lot); each worker makes exactly one call; workers run concurrently; the orchestrator does not re-plan. Real workers usually run a tool-using loop of their own, which is where Anthropic's "about 15×" comes from. So the arithmetic says something narrower than "orchestrators are cheap": orchestration itself adds little; what costs is many workers each doing real work, and what you buy is time and breadth.
The research behind orchestrators
The pattern appeared in research in 2023. HuggingGPT (Shen et al., Zhejiang University and Microsoft Research Asia, March 2023) used a language model as a controller over expert models from the Hugging Face hub in four stages: task planning, model selection, task execution and response generation. AutoGen (Wu et al., Microsoft Research and three universities, August 2023) made multi-agent conversation a programming framework. Magentic-One (Fourney et al., Microsoft Research, November 2024) is the most carefully engineered of the three, and its orchestrator keeps two explicit records.
Research notes: HuggingGPT, Magentic-One's ledgers and results, AutoGen, and Anthropic's prompting rules (optional reading)
Orchestrator-workers in practice
| Fixed sectioning | Orchestrator-workers | |
|---|---|---|
| Pros | predictable cost and latency; each section testable | handles tasks whose split depends on the input or on early findings; breadth in parallel |
| Cons | cannot adapt the split | cost is the sum of every worker's loop (Anthropic reports about 15 times a chat); plans and briefs can be wrong; harder to test |
| When to pick it | the parts are known in advance | open-ended, breadth-first, high-value tasks with independent sub-questions |
| Production failure | a part nobody planned for | duplicated or contradictory workers, runaway spawning, a synthesis that drops findings |
What to measure before you decide
| Measurement | How | What it tells you |
|---|---|---|
| Plan quality | grade a sample of plans before execution | whether the orchestrator decomposes well, separately from execution |
| Worker overlap | compare workers' queries or sources per task | duplicated work from vague briefs |
| Workers and tokens per task | from traces, against task difficulty | whether effort scales with the question (Anthropic's early failure) |
| Synthesis fidelity | check that each worker's key finding appears in the answer | lost or contradicted findings |
| Value per task | what a good answer is worth against its token cost | whether the pattern pays at all |
Discussion
- Is orchestrator-workers a workflow or an agent? The orchestrator decides the split, so it is the most agent-like of the five. In practice the outer frame (plan, fan out, synthesise, stop) is code and the inside is the model's, which is where the testing burden lands.
- Should workers share context? Isolation keeps contexts small and stops one worker's error spreading; sharing lets workers build on each other. My view: isolate by default and pass findings explicitly through the orchestrator, because that path is traceable.
- When is the 15-times cost worth it? When a wrong or slow answer is expensive and the question is breadth-first. A monthly operations question at Northwind qualifies; answering each support ticket does not.
4.6 Evaluator-optimiser: generate, check, refine, stop
The example: rewriting Northwind's weekly update
Every week Northwind's team leads write paragraphs about their work, and the internal-communications editor wants them rewritten for the all-staff email under strict rules: 55 to 75 words, exactly four sentences, no sentence over 20 words, three named keywords, and none of ten banned words ("very", "basically", "leverage" and so on). A single call often misses one rule. The evaluator-optimiser pattern adds a second role: something that checks the draft and sends back what is wrong, in a loop that stops by a rule.
The arithmetic of the stop, with an imperfect evaluator
If each round passes with probability and the evaluator is perfect, the chance of passing within rounds is and the cost per accepted draft is one round's cost divided by , whatever is (Chapter 3 derived this). Assumptions: rounds are independent with the same , each round costs the same, and the evaluator never errs. Feedback-driven rounds break the first assumption (feedback usually raises after the first round), so this is a floor on the benefit, not a forecast.
The evaluator is never perfect. Write for the chance it accepts a good draft and for the chance it accepts a bad one (its false-accept rate). The share of drafts the loop stops on that are actually good is, by Bayes' rule,
the same formula as the cascade's check in Section 4.3. With , and it gives 97.8%; with and the same evaluator, 68%. The harder the task, the more the evaluator's false-accept rate matters. And a low (rejecting good drafts) has its own cost: more rounds, and rewrites of drafts that were already fine.
Our experiment: a code checker against a judge model
ch4_evaluator.py rewrites eight source paragraphs under the five rules, in two loops from the same start, each allowed up to four rounds. In the programmatic loop a Python checker lists the rules each draft breaks, with measured values ("C3: a sentence of 24 words, limit 20"), and the loop stops when the list is empty. In the judge loop a separate model call marks each rule pass or fail, its verdict is fed back, and the loop stops when it says everything passes. Every draft in both loops is also scored by the other evaluator, so we can count how often the judge and the code disagree. Because the programmatic loop is graded by the same code that gives it feedback, we add a held-out check that no loop ever sees: two key facts per source paragraph (for example "staging" and "audit log" in the incident report) must survive the rewrite. The run cost about two cents.
# code/agents/ch4_evaluator.py (abridged): one loop; arm decides who the critic is. The held-out fact check is applied once, at the end.
for rnd in range(1, MAX_ROUNDS + 1):
text, ug = llm.ask(msgs, max_tokens=400) # generator
viol = check(text, keywords) # programmatic checker: exact counts, never wrong about counting
jfails, notes, uj = judge(text, keywords) # judge model: pass/fail per constraint
critic_fails = viol if arm == 'programmatic' else jfails # who decides whether to stop
if not critic_fails:
break
feedback = viol if arm == 'programmatic' else [f'{c}: the reviewer marked this constraint as failed. {notes}' for c in jfails]
msgs += [{'role': 'assistant', 'content': text}, {'role': 'user', 'content': REGEN.format(viol='\n- '.join(feedback))}]The judge cannot count. Over 40 drafts it agreed with the checker's overall verdict 40% of the time, with 2 false passes and 22 false fails. It was worst on the word count (32% agreement) and on the sentence-length rule (62%), the two rules that require counting, and best on keywords (90%), which require only looking.
False fails are not harmless. In this run the judge mostly rejected drafts that were fine. Its loop averaged 3.1 rounds against 1.9, five of its eight paragraphs hit the four-round cap, and with the judge's own call in every round it cost $0.0005 per round against $0.0002. Worse, rewriting a passing draft can damage it: on paragraph 5 the first draft met every rule, the judge rejected it three times, and the final draft had dropped one of the held-out facts (the hourly queries that drive the warehouse costs). In a development run the judge's errors leaned the other way, towards false passes; the direction varies, the unreliability does not.
A false pass ships a broken draft. On paragraph 1 the judge accepted a 52-word draft against a 55-word minimum after one round. The programmatic loop, whose checker cannot miscount, stopped on 7 of 8 paragraphs, and all 7 truly passed and kept their held-out facts. Its one failure was a paragraph about app crashes where the model kept writing one 24-word sentence across four rounds, the "no progress" case for which a stop rule should hand the draft to a person.
The research behind judges
The reference study of model judges is Lianmin Zheng and colleagues' "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (the LMSYS group at Berkeley and others, June 2023). It found strong judges agreeing with human preferences about as often as humans agree with each other, "over 80% agreement", and it catalogued the biases.
G-Eval (Liu et al., Microsoft, March 2023) showed how to make a judge more consistent: have the model write evaluation steps from the criteria, fill in a form, and weight the score by the probabilities of each grade. Its Spearman correlation with human ratings of summaries was 0.514 with GPT-4, ahead of every earlier automatic metric; the authors also warn that model evaluators may prefer model-written text. Chapter 3's caution applies here too: Huang et al. (2023) found that asking a model to correct its own reasoning without outside feedback often made it worse. A judge is outside feedback only to the extent that it sees something the generator did not, or is better than the generator at the specific check.
Research notes: the judge paper's abstract, G-Eval, and Anthropic's conditions for the pattern (optional reading)
Evaluator-optimiser in practice
| Code evaluator (tests, schema, counts) | Model judge | |
|---|---|---|
| Pros | exact, free, fast; never miscounts; specific feedback | judges what code cannot: tone, helpfulness, faithfulness to a source |
| Cons | only checks what you can write down | false passes and false fails; position and verbosity bias; one extra call per round |
| When to pick it | any rule you can express in code | qualities with no programmatic test, after measuring its and |
| Production failure | a check that tests the wrong thing, passed reliably | a judge that drifts with a model update and nobody re-measures it |
What to measure before you decide
| Measurement | How | What it tells you |
|---|---|---|
| Judge and | run the judge on a few hundred labelled drafts | the precision of the stop, via Bayes |
| Position and verbosity sensitivity | swap order; pad an answer without adding content | whether the judge's verdicts survive changes that should not matter |
| Rounds to pass, and the no-progress rate | from traces | where to cap rounds; how often to hand off to a person |
| Damage from extra rounds | held-out checks on drafts that passed early | whether false fails are degrading good work |
| Cost per round and per accepted draft | API usage per round | the price of each extra round |
Discussion
- Is a judge outside feedback? Partly. It sees the draft fresh, without the generator's reasoning, but it shares the model's blind spots if it is the same model. Using a different model as judge reduces shared blind spots; it does not remove the need to measure the judge.
- What should the loop optimise? Whatever the evaluator checks, exactly as our gated chain did. If the evaluator is narrow, add held-out checks the loop never sees, as we did with the key facts.
- How many rounds? Our programmatic loop passed most paragraphs by round two; most of the judge loop's extra rounds were spent rewriting drafts that already passed. A cap of two to four rounds, with a no-progress stop, is a reasonable start, then adjust from traces.
4.7 Composing patterns: the workflow as a state machine
The example: Northwind's support workflow, all together
Real systems combine the patterns. Northwind's support workflow, built in full in ch4_workflow.py, puts a router in front (code rules first, a model call only when the rules do not decide), a gated chain behind it (extract, then draft, each schema-validated with one retry), a policy check in code, a human approval step for refunds over £50, and an idempotent send. Around those nodes sits the plumbing every production workflow needs: typed state, budgets, checkpoints, resume, and a trace. A larger version could hang an orchestrator off one branch (for "why are complaints up?" questions) and an evaluator loop before the send; the plumbing would not change.
The complete program
The core of the runner is a loop over the current node, with a checkpoint after every node:
# code/agents/ch4_workflow.py (abridged): the runner. Every node updates the typed State and names the next node.
def run(st, crash_after=None):
try:
while st.node not in ('done', 'needs_human', 'awaiting_approval'):
t0 = time.time()
if st.node == 'extract':
obj, use, retried = extract(st) # schema-validated (pydantic), one retry with the errors
if obj is None:
span(st, 'extract', 'gate failed twice -> person', t0, use); st.node = 'needs_human'
else:
st.facts = obj.model_dump(); span(st, 'extract', 'ok', t0, use); st.node = 'draft'
elif st.node == 'review':
amt = st.facts.get('refund_amount_gbp') or 0
if amt > REFUND_LIMIT and st.approved is None:
span(st, 'review', 'PAUSE for approval', t0); st.node = 'awaiting_approval' # human-in-the-loop interrupt
else:
span(st, 'review', 'policy ok', t0); st.node = 'send'
... # route, faq, draft and send follow the same shape
checkpoint(st) # state saved after every node: resume starts here
except BudgetExceeded as e: # at most 6 model calls and 8,000 tokens per ticket
span(st, st.node, f'STOP: {e}', time.time()); st.node = 'needs_human'; checkpoint(st)
return stThe demonstration runs six tickets, kills the process on purpose after ticket T3's draft is checkpointed, resumes every unfinished ticket from its checkpoint, has a person approve the paused refund, and then replays a send to show that it cannot happen twice. It made 12 model calls and cost $0.0015.
Four behaviours in that output are the reasons the plumbing exists. Resume skips finished work: T3 continued at review with its two earlier calls already counted, and extract and draft were not run again. The interrupt waits: T5's £649 refund stopped at awaiting_approval with its state on disk, and continued from there after approval. The side effect is idempotent: replaying T1's send found its key in the outbox and did nothing. The router used the model only when needed: four tickets were routed by keyword rules at no cost; the Spanish ticket (T6) and the shipping question (T4) fell through the rules and cost one small call each.
The full program: code/agents/ch4_workflow.py (about 260 lines; optional reading)
"""Chapter 4, Section 4.7: one complete, runnable composed workflow for Northwind Home's support tickets, with the real API
(gpt-5-mini, reasoning_effort=minimal). Everything the chapter recommends, in about 250 lines:
typed state (a dataclass) and typed step outputs (pydantic schemas, validated in code);
a router that is code first and asks the model only when the rules do not decide;
a gated chain: extract -> draft, each step schema-validated, retried once with the exact errors, then handed to a person;
a policy check in code (word limit, banned phrases, order id quoted) after the draft;
a human-in-the-loop interrupt: refunds over GBP 50 pause the run until someone approves;
an idempotent side effect: replies go to an outbox keyed by an idempotency key, so a resumed run cannot send twice;
budgets per ticket (model calls, tokens) that stop the run instead of letting it spin;
a checkpoint after every node, and resume from the last checkpoint after a crash;
a trace: one JSON line per node (trace id, node, outcome, tokens, milliseconds) in results/ch4_workflow_trace.jsonl.
The demo processes six tickets, injects one crash after the draft step of ticket T3, resumes it, and approves the paused refund."""
import hashlib, json, os, re, time, uuid
from dataclasses import dataclass, field, asdict
from typing import Literal, Optional
from pydantic import BaseModel, Field, ValidationError
from common import Log, save, RESULTS
import ch4_llm as llm
log = Log('ch4_workflow')
llm.set_log(log)
CKPT = os.path.join(RESULTS, 'ch4_workflow_ckpt')
OUTBOX = os.path.join(RESULTS, 'ch4_workflow_outbox.json')
TRACE = os.path.join(RESULTS, 'ch4_workflow_trace.jsonl')
MAX_CALLS, MAX_TOKENS, REFUND_LIMIT = 6, 8000, 50.0
BANNED = ['24 hours', 'immediately', 'guarantee', 'today', 'right away', 'as soon as possible', 'shortly']
# ---------------------------------------------------------------- typed step outputs (the schemas the gates enforce)
class Facts(BaseModel):
issue_type: Literal['billing', 'delivery', 'damaged_or_faulty', 'account', 'other']
order_id: Optional[str] = Field(default=None, pattern=r'^ORD-\d{5}$')
refund_amount_gbp: Optional[float] = Field(default=None, ge=0)
class Draft(BaseModel):
reply: str = Field(min_length=20)
# ---------------------------------------------------------------- typed workflow state
@dataclass
class State:
ticket_id: str
text: str
trace_id: str = field(default_factory=lambda: uuid.uuid4().hex[:8])
node: str = 'route' # the next node to run; 'done' and 'needs_human' and 'awaiting_approval' are terminal or paused
route: Optional[str] = None
facts: Optional[dict] = None
reply: Optional[str] = None
approved: Optional[bool] = None
calls: int = 0
tokens: int = 0
notes: list = field(default_factory=list)
class BudgetExceeded(Exception):
pass
class InjectedCrash(Exception):
pass
# ---------------------------------------------------------------- infrastructure: checkpoints, trace, model calls with budgets
def checkpoint(st):
os.makedirs(CKPT, exist_ok=True)
json.dump(asdict(st), open(os.path.join(CKPT, st.ticket_id + '.json'), 'w'), indent=1)
def load(ticket_id):
p = os.path.join(CKPT, ticket_id + '.json')
return State(**json.load(open(p))) if os.path.exists(p) else None
def span(st, node, outcome, t0, use=None):
rec = {'trace_id': st.trace_id, 'ticket': st.ticket_id, 'node': node, 'outcome': outcome, 'ms': round((time.time() - t0) * 1000),
'tokens': (use or {}).get('input', 0) + (use or {}).get('output', 0)}
open(TRACE, 'a').write(json.dumps(rec) + '\n')
log(f' [{st.trace_id}] {st.ticket_id} {node:<9} {outcome:<34} {rec["tokens"]:>5} tok {rec["ms"]:>6} ms')
def call(st, prompt):
if st.calls >= MAX_CALLS or st.tokens >= MAX_TOKENS:
raise BudgetExceeded(f'budget: {st.calls} calls, {st.tokens} tokens')
text, use = llm.ask(prompt, max_tokens=600)
st.calls += 1
st.tokens += use['input'] + use['output']
return text, use
def validated(st, prompt, model_cls, extra_check=None):
"""A gated step: call, parse, validate against the schema (and an optional extra check), retry once with the exact errors."""
total = {'input': 0, 'output': 0}
for attempt in range(2):
text, use = call(st, prompt)
for k in total:
total[k] += use[k]
try:
obj = model_cls.model_validate(llm.parse_json(text) or {})
problems = extra_check(obj) if extra_check else []
except ValidationError as e:
obj, problems = None, [f'{".".join(str(x) for x in err["loc"])}: {err["msg"]}' for err in e.errors()]
if not problems:
return obj, total, attempt
prompt = prompt + '\n\nYour previous answer failed these checks:\n- ' + '\n- '.join(problems) + '\nReply with a corrected JSON object only.'
return None, total, 1
# ---------------------------------------------------------------- the nodes
def route(st):
t = st.text.lower()
if re.search(r'\b(refund|charged|invoice|broken|cracked|damaged|parcel|delivery|order)\b', t): # code decides when it can
return 'support', None
text, use = call(st, 'Is this message a support request about an order or account (answer "support") or a general question '
f'about products, shipping zones or opening hours (answer "faq")? Reply with one word.\n\nMessage: {st.text}')
return ('faq' if 'faq' in text.lower() else 'support'), use
def extract(st):
prompt = ('Extract facts from this support ticket. Reply with ONE JSON object with keys: "issue_type" (one of billing, delivery, '
'damaged_or_faulty, account, other), "order_id" (the order number as ORD-NNNNN, five digits, or null), '
f'"refund_amount_gbp" (a number if the customer asks for money back and states an amount, else null).\n\nTicket: {st.text}')
return validated(st, prompt, Facts)
def policy_problems(reply, order_id):
p = []
if len(reply.split()) > 80:
p.append(f'the reply has {len(reply.split())} words; the limit is 80')
hits = [b for b in BANNED if b in reply.lower()]
if hits:
p.append(f'the reply promises timing: {hits}; remove these phrases')
if order_id and order_id not in reply:
p.append(f'the reply must quote the order id {order_id}')
return p
def draft(st):
f = st.facts
prompt = ('Write a short, polite reply (at most 80 words) to this customer, for Northwind Home. Do not promise any timing. '
+ (f'Quote the order id {f["order_id"]}. ' if f.get('order_id') else '')
+ ('Say that the refund request has been passed for approval. ' if f.get('refund_amount_gbp') else '')
+ 'Reply with ONE JSON object: ' + json.dumps({'reply': '...'}) + f'\n\nTicket: {st.text}\nFacts: {json.dumps(f)}')
return validated(st, prompt, Draft, lambda d: policy_problems(d.reply, f.get('order_id')))
def send(st):
"""The side effect. Idempotent: the key is derived from the ticket and the reply, and the outbox ignores a key it has seen."""
key = f'{st.ticket_id}:' + hashlib.sha256(st.reply.encode()).hexdigest()[:12] # stable across processes, unlike hash()
box = json.load(open(OUTBOX)) if os.path.exists(OUTBOX) else {}
if key in box:
return 'already sent (idempotency key seen)'
box[key] = {'ticket': st.ticket_id, 'reply': st.reply, 'at': time.strftime('%H:%M:%S')}
json.dump(box, open(OUTBOX, 'w'), indent=1)
return 'sent'
# ---------------------------------------------------------------- the graph runner
def run(st, crash_after=None):
"""Run from st.node until the workflow finishes, pauses or fails; checkpoint after every node."""
try:
while st.node not in ('done', 'needs_human', 'awaiting_approval'):
t0 = time.time()
if st.node == 'route':
st.route, use = route(st)
span(st, 'route', f'-> {st.route}' + (' (model)' if use else ' (rule)'), t0, use)
st.node = 'extract' if st.route == 'support' else 'faq'
elif st.node == 'faq':
st.notes.append('sent to the FAQ answerer (out of scope for this demo)')
span(st, 'faq', 'handed to the FAQ workflow', t0)
st.node = 'done'
elif st.node == 'extract':
obj, use, retried = extract(st)
if obj is None:
span(st, 'extract', 'gate failed twice -> person', t0, use); st.node = 'needs_human'
else:
st.facts = obj.model_dump()
span(st, 'extract', f'ok{" after retry" if retried else ""} {st.facts["issue_type"]} {st.facts["order_id"]}', t0, use)
st.node = 'draft'
elif st.node == 'draft':
obj, use, retried = draft(st)
if obj is None:
span(st, 'draft', 'gate failed twice -> person', t0, use); st.node = 'needs_human'
else:
st.reply = obj.reply
span(st, 'draft', f'ok{" after retry" if retried else ""} ({len(st.reply.split())} words)', t0, use)
st.node = 'review'
elif st.node == 'review':
amt = st.facts.get('refund_amount_gbp') or 0
if amt > REFUND_LIMIT and st.approved is None:
span(st, 'review', f'refund GBP {amt:.2f} > {REFUND_LIMIT:.0f}: PAUSE', t0); st.node = 'awaiting_approval'
elif st.approved is False:
span(st, 'review', 'refund rejected by a person', t0); st.node = 'needs_human'
else:
span(st, 'review', 'policy ok' + (' (approved)' if st.approved else ''), t0); st.node = 'send'
elif st.node == 'send':
span(st, 'send', send(st), t0); st.node = 'done'
checkpoint(st)
if crash_after and st.node == crash_after[1] and st.ticket_id == crash_after[0]:
raise InjectedCrash(f'process killed after checkpointing, before running {st.node}')
except BudgetExceeded as e:
span(st, st.node, f'STOP: {e}', time.time()); st.node = 'needs_human'; checkpoint(st)
return st
TICKETS = [('T1', 'Order #48213 arrived with the screen cracked. Please send a replacement.'),
('T2', 'I was charged twice for ORD-10988, please refund the extra GBP 39.99.'),
('T3', 'Where is my parcel? Tracking has not moved in 6 days. Order 77120.'),
('T4', 'Do you ship to the Isle of Man?'),
('T5', 'The sofa from order 61234 has a torn seam. I want my money back, GBP 649.'),
('T6', 'Mi pedido 33310 llego roto.')]
if __name__ == '__main__':
import shutil
for p in [CKPT]:
shutil.rmtree(p, ignore_errors=True)
for p in [OUTBOX, TRACE]:
if os.path.exists(p):
os.remove(p)
log(f'== A composed workflow for Northwind Home support ({llm.MODEL}, minimal effort): router -> gated chain -> policy -> approval -> send ==')
log(f'budgets per ticket: {MAX_CALLS} model calls, {MAX_TOKENS} tokens; refunds over GBP {REFUND_LIMIT:.0f} need a person')
log('')
log('-- first run (a crash is injected in T3 after the draft is checkpointed) --')
for tid, text in TICKETS:
st = State(tid, text)
try:
run(st, crash_after=('T3', 'review'))
except InjectedCrash as e:
log(f' !! {tid}: {e}')
log('')
log('-- resume: load every checkpoint that is not finished and continue from its saved node --')
for tid, _ in TICKETS:
st = load(tid)
if st.node not in ('done', 'needs_human', 'awaiting_approval'):
log(f' {tid}: resuming at node "{st.node}" (calls so far {st.calls}; extract and draft are NOT re-run)')
run(st)
log('')
log('-- a person approves the paused refund, and the run continues from the checkpoint --')
for tid, _ in TICKETS:
st = load(tid)
if st.node == 'awaiting_approval':
st.approved = True
st.node = 'review'
log(f' {tid}: approved by a person')
run(st)
log('')
log('-- replaying "send" for T1 to show idempotency --')
st = load('T1'); t0 = time.time(); span(st, 'send', send(st), t0)
log('')
final = [load(t) for t, _ in TICKETS]
log(f'{"ticket":<7} {"route":<8} {"final node":<12} {"calls":>5} {"tokens":>7} facts')
for st in final:
f = st.facts or {}
log(f'{st.ticket_id:<7} {str(st.route):<8} {st.node:<12} {st.calls:>5} {st.tokens:>7} {f.get("issue_type", "-")} {f.get("order_id", "-")} '
f'{("refund " + str(f.get("refund_amount_gbp"))) if f.get("refund_amount_gbp") else ""}')
box = json.load(open(OUTBOX))
spans = [json.loads(l) for l in open(TRACE)]
U = llm.snapshot()
log(f'outbox: {len(box)} replies sent for {sum(1 for s in final if s.node == "done" and s.route == "support")} finished support tickets; '
f'trace: {len(spans)} spans in results/ch4_workflow_trace.jsonl')
log(f'total spend: ${llm.cost(U):.4f} ({U["calls"]} calls); stub used: ' + ('YES' if llm.stub_used() else 'no'))
save('ch4_workflow', {'states': [asdict(s) for s in final], 'spans': spans, 'outbox': box, 'usage': U, 'stub': llm.stub_used()})Research notes: LangGraph's interrupt (optional reading)
Observability to design in now
Every node in the program writes one line to a trace file: the trace id of the run, the ticket, the node, the outcome, the tokens and the milliseconds. That is the minimum. With it you can answer the questions this chapter has asked of every pattern: which step fails first, how often each gate retries, how many tickets the router sends to the model, how long approvals wait, what each ticket costs. Chapter 8 builds this into proper tracing with spans, parent ids and dashboards; the decision to make now is that every node emits a record with the same trace id.
Composition in practice
| Plumbing | What it prevents | How the program does it |
|---|---|---|
| Typed state and schemas | a step reading a field that is missing or malformed | a dataclass for state; pydantic models for step outputs |
| Gates with one retry | format failures flowing downstream | validated(): parse, validate, retry once with the errors |
| Budgets per run | a loop or retry storm burning money | a call and token cap that stops the run and hands it to a person |
| Checkpoints and resume | a crash losing work, or redoing paid calls | state written after every node; load() continues from node |
| Idempotent side effects | duplicate emails or refunds after a resume | an outbox keyed by a stable hash of ticket and reply |
| Human interrupt | money moving without approval | awaiting_approval state, resumed after a person decides |
| Trace | not knowing where or why runs fail | one JSON line per node with the trace id |
Discussion
- Framework or hand-written? Our runner is about thirty lines; LangGraph, Temporal and the vendor SDKs add persistence, retries, visualisation and hosting. A framework pays for itself when you need durable waits of hours or days and many workflows; the concepts are the same either way.
- Where should budgets live? In code, per run, enforced before each call, as in the program. A budget written only in a prompt is a suggestion.
- What is the first thing to break when the workflow grows? Usually the interfaces: a schema changes in one step and the next step still expects the old one. Version schemas and put the version in the trace.
4.8 Workflows in production: what companies built and what they report
This section reads engineering write-ups for the same four things each time: which pattern, why, what went wrong, and what they report. Every quotation was checked against the live page. Each case separates what the company reports from our reading of it, and attributes numbers to the exact system they describe. Where a page gives no number, none is given here. Pages that block automated browsers are quoted with a link rather than screenshotted.
Uber: a chain of filters, and a router of intents
Anthropic: orchestrator-workers for research
Section 4.5 read the post in detail. What Anthropic reports: a lead agent with parallel subagents "outperformed single-agent Claude Opus 4 by 90.2%" on an internal research eval; multi-agent systems "use about 15× more tokens than chats"; early versions spawned "50 subagents for simple queries"; parallelism "cut research time by up to 90% for complex queries"; the lead agent waits for each set of subagents, which "creates bottlenecks". For evaluation they use a single model judge with a rubric (factual accuracy, citation accuracy, completeness, source quality, tool efficiency) scoring 0.0 to 1.0 with a pass-fail grade, plus human testing, starting from "a set of about 20 queries". Our reading: the pattern paid where subtasks were independent and answers valuable; the failures were orchestration failures (effort, briefs, stopping) fixed in the orchestrator's prompt and in code limits.
DoorDash: a guarded pipeline for support
DoorDash's September 2024 post "Path to high-quality LLM-based Dasher support automation" (DoorDash's site blocks automated browsers, so it is quoted here) describes a support system for delivery drivers. What DoorDash reports: a fixed retrieval pipeline with a two-tier guardrail (a cheap semantic-similarity check, then a model-based evaluator) that can retry or hand the conversation to a person, plus an offline model judge on five quality dimensions. "This guardrail system has successfully reduced overall hallucinations by 90% and cut down potentially severe compliance issues by 99%"; the guardrail's latency "is a notable drawback", and a more sophisticated guardrail model was dropped because "increased response times and heavy usage of model tokens made it prohibitively expensive". A November 2025 post from the same company ("Beyond Single Agents") argues for starting with "deterministic workflows" and warns that "you can't jump straight to sophisticated, multi-agent collaboration". Our reading: a chain with an evaluator at the end, ordered cheap check first and model check second, with a human fallback rather than a long retry loop; the cost of the evaluator shaped the design.
More production cases: Stripe, Airbnb, LinkedIn, Uber Genie, Spotify, Microsoft and Amazon (optional reading)
Stripe and Airbnb: workflows written as state machines. Two coding pipelines describe themselves in this chapter's vocabulary. Stripe (February 2026, "Minions") reports that its background coding agents follow "blueprints", workflows "defined in code" that look like "a state machine that intermixes deterministic code nodes and free-flowing agent nodes", with at most two rounds of continuous integration "since CI runs cost tokens, compute, and time". Airbnb (March 2025, "Accelerating Large-Scale Test Migration with LLMs", quoted because the page blocks automated browsers) migrated about 3,500 test files with a per-file pipeline "modeled ... like a state machine", retrying steps with the validation errors "until they passed or we reached a limit"; it reports 75% of files migrated "in just four hours", 97% after four days of tuning, and the remaining 3% finished by hand. Our reading: both are gated chains with code evaluators (tests, linters, CI) and capped retries, with a model inside some nodes and people at the end of the tail.
Spotify, Honk (December 2025). Spotify's background coding agent ends with deterministic verifiers (build, tests, formatting) and then a model judge that compares the diff with the original prompt: "The judge is simple. It uses the diff of the proposed change and the original prompt, and sends them to an LLM for evaluation." The post reports that "the judge vetoes about a quarter" of thousands of sessions and that "the agent is able to course correct half the time", and states plainly: "We have yet to invest in evals for our judge." Our reading: an evaluator placed after the code checks, in the order this chapter recommends, whose own and are not yet measured.
LinkedIn, SQL Bot (December 2024) checks generated queries with validators that "access new information not available to the query writer" (tables and fields exist; EXPLAIN runs) and feeds errors to a correction step (Chapter 3 reads it in detail). Microsoft's Magentic-One (Section 4.5) is the research orchestrator with explicit task and progress ledgers. Amazon's code-transformation agent in Q Developer has developers "review and iterate on the plan before the agent implements it", then builds and tests the result (Chapter 3). We could not verify public engineering write-ups with workflow details for Netflix, Instacart or Klarna, so they are not included.
The table
| Company, system | Pattern | Why (as stated) | Reported problem or limit | Reported result |
|---|---|---|---|---|
| Uber, uReview | sectioned generation, then a grader and filters | single prompts gave false positives and low-value comments | noise; thresholds need per-language tuning | 75% of comments marked useful; 65% addressed |
| Uber, QueryGPT | intent router, then a chain; human edits the tables | accuracy fell as tables grew | hallucinated tables and columns; ~5% run-to-run variance | not given as a single number |
| Uber, Genie (EAg-RAG) | chain of pre- and post-processing agents; offline judge | incomplete or wrong answers; slow expert evaluation | evaluation took experts weeks | +27% relative acceptable answers, -60% relative incorrect advice |
| LinkedIn, Premium assistant | router, retrieval, generation | different difficulty per step | ~10% of structured outputs invalid; last 15% of quality slow | 80% of target in one month; four more months towards 95% |
| DoorDash, Dasher support | pipeline with a two-tier guardrail and human fallback | hallucination and compliance risk | guardrail latency; a stronger guardrail too costly | -90% hallucinations, -99% severe compliance issues |
| Spotify, Honk | agent with code verifiers, then a model judge | agents straying outside the prompt | judge not yet evaluated | judge vetoes about a quarter; half recover |
| Stripe, Minions | state machine of code nodes and agent nodes; CI as evaluator | CI rounds cost tokens and time | diminishing returns from more CI rounds | capped at two CI rounds |
| Airbnb, test migration | per-file state machine with capped retries | a known transformation over 3,500 files | a long tail automation could not fix | 75% in four hours; 97% in four days; rest by hand |
| Anthropic, Research | orchestrator-workers with parallel subagents | breadth-first questions | 15x chat tokens; over-spawning; synchronous bottleneck | +90.2% on an internal eval |
Three things hold across the table, as patterns rather than laws. Most systems are chains, with a router in front or an evaluator behind; the orchestrator appears where the question is open-ended and valuable. Evaluators are usually placed after cheap code checks, and their cost and latency shaped the designs (DoorDash dropped a stronger guardrail; Stripe capped CI rounds). And the teams that report the most progress also report how they measured each step, not only the end result.
Discussion
- Why are most production systems chains? The tasks automated first are the ones whose steps are known, and a known chain is cheaper to run, test and explain. That is a reasonable order of work, not a lack of ambition.
- Which reported numbers can you compare? Few. Each company measures its own system on its own data with its own definition of success (useful comments, acceptable answers, vetoed sessions). Read them as evidence that the pattern was worth keeping, not as benchmarks.
- What is missing from the write-ups? Mostly the evaluators' own error rates: how often the grader, guardrail or judge is wrong. Spotify says so openly; the rest are silent. That is the measurement to add first in your own system.
4.9 Choosing and evolving a workflow
The decision is not a single choice but an order of additions, each justified by a measured problem. The questions below map task properties to patterns.
| Task property | How to measure it | What it argues for |
|---|---|---|
| Variability of inputs | cluster a sample of real inputs; count the kinds | one kind: a chain; several kinds with different handling: a router in front |
| Independence of subtasks | for each pair of subtasks, does one need the other's output? | independent: parallel sections; dependent: a chain; unknowable in advance: an orchestrator |
| Verifiability | list the checks you can write in code, and the ones that need judgement | code checks: gates and an evaluator loop; judgement only: a measured judge, or a person |
| Latency budget | the product's p95 target against the chain's sequential depth | tight: parallelise, use smaller models per step, cut steps |
| Cost per task | tokens per run from traces, times traffic | high volume: routers and cascades; high value per task: orchestrators may pay |
| Blast radius | what happens if a wrong output ships: a reply, a refund, a deleted record | high: a human interrupt before the side effect, and idempotent steps |
The migration path in the figure is a heuristic: the patterns compose, so a system may need an evaluator before it needs a router, and some tasks start as orchestrators. What does hold is the order of evidence. Each addition should answer a failure you have seen in traces and labelled data, and each should be re-measured on accuracy, cost per correct answer and latency once it is in.
The one-rule summary of the chapter: hand the model a decision only when you cannot write it in code, and measure every decision you hand over. A chain hands over none; a router hands over one, measurable with a confusion matrix; parallel calls hand over none and buy time; an orchestrator hands over the shape of the work, measurable by grading plans and counting duplicated effort; an evaluator hands over the decision to stop, measurable by its false-accept and false-reject rates. The experiments in this chapter found the same thing from four directions: the model was good at the content of each step and unreliable at the decisions around it (whether its answer was good enough, whether a draft had 55 words, whether a question was hard), and code made those decisions better wherever code could make them.
Discussion
- Where should Northwind start? With the gated chain of Section 4.2, a labelled set of a few hundred tickets, and per-step evals. Everything else waits for a measured reason.
- When is it right to skip a step on the path? When the task's structure is obvious: code generation against tests goes straight to an evaluator loop with the tests as the evaluator; a research question goes straight to an orchestrator. Skipping is fine if you still take the measurement that would have justified the step.
- What should be re-measured when the model is upgraded? Every model decision: the router's confusion matrix, the judge's error rates, the per-step success rates. A new model can improve the content and change the decisions in either direction.
Exercises
- Northwind's chain has three steps at 95%, 92% and 98% per step on a labelled set. Compute the end-to-end success under the independence assumption, say which step a gate would help most, and describe one way the independence assumption could fail on real tickets.
- A cascade's check accepts 90% of right cheap answers and 25% of wrong ones. Compute the precision of accepted answers when the cheap model is right 80% of the time and when it is right 40% of the time, and decide in which case you would ship the cascade.
- Design the brief an orchestrator should give one worker for "why did delivery complaints rise last quarter?": objective, output format, tools and sources, boundaries. Then list two measurements that would tell you whether workers are duplicating each other.
- Your evaluator loop uses a model judge for tone and a word count in code. Traces show the loop runs four rounds on 30% of drafts. Describe how you would find out whether the judge is producing false fails, and what you would change if it is.
- In
ch4_workflow.py, the send step is idempotent because of its outbox key. Name two other side effects a support workflow might have, and how you would make each idempotent.
Key takeaways
- A workflow is model calls arranged by control flow the engineer wrote; an agent lets the model choose the next step. Most production systems are workflows with a model inside some nodes.
- Draw the system as a graph of nodes and edges, and mark which edges code decides and which the model decides. Each decision handed to the model needs its own measurement.
- Prompt chains make each step easy and testable. End-to-end success is roughly the product of per-step success, so gates between steps matter: in our run, gates with one retry lifted a loosely specified chain from 12 to 16 of 20, and could not catch valid-but-wrong values.
- Routers are judged by their confusion matrix, not their accuracy. Our model router sent one of 24 questions to the strong model and caught none of the seven that needed it; a cascade reached 88% at about two-thirds of the strong model's cost per correct answer.
- The false-accept rate of a check is not one minus its precision: precision depends on the base rate, .
- Parallel calls buy time, not tokens (4.5 times faster in our run). A plurality vote needs the right answer to be the most frequent, not a majority; correlated errors cut the gain, and our vote gained nothing. Best-of-n with a verifier helped, by less than independence predicts.
- Orchestrator-workers fits open-ended, breadth-first, valuable tasks. Anthropic reports about 15 times the tokens of a chat; the failures they report are about briefs, effort and stopping.
- In an evaluator-optimiser loop, put code checks first and measure any judge before it stops a loop. Our judge agreed with a word count 32% of the time, and its false fails both cost rounds and damaged a good draft.
- Production plumbing is part of the pattern: typed state, schema checks, budgets, checkpoints, idempotent side effects, human interrupts and one trace line per node, all shown working in
ch4_workflow.py. - Hand the model a decision only when you cannot write it in code, and re-measure every handed-over decision after each change.
References
Papers
- Tongshuang Wu, Michael Terry, Carrie J. Cai. AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. CHI 2022 (arXiv October 2021).
- Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, Ed Chi. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. ICLR 2023 (arXiv May 2022).
- Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, Ashish Sabharwal. Decomposed Prompting: A Modular Approach for Solving Complex Tasks. ICLR 2023 (arXiv October 2022).
- Lingjiao Chen, Matei Zaharia, James Zou. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv May 2023.
- Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, Ahmed Hassan Awadallah. Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing. ICLR 2024 (arXiv April 2024).
- Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, Ion Stoica. RouteLLM: Learning to Route LLMs with Preference Data. ICLR 2025 (arXiv June 2024).
- Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, Deheng Ye. More Agents Is All You Need. TMLR 2024 (arXiv February 2024).
- Xinyun Chen, Renat Aksitov, Uri Alon, Jie Ren, Kefan Xiao, Pengcheng Yin, Sushant Prakash, Charles Sutton, Xuezhi Wang, Denny Zhou. Universal Self-Consistency for Large Language Model Generation. arXiv November 2023.
- Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, Azalia Mirhoseini. Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. arXiv July 2024.
- Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, Yueting Zhuang. HuggingGPT (full title on arXiv). NeurIPS 2023 (arXiv March 2023).
- Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, Chi Wang. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv August 2023.
- Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Erkang Zhu, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, Peter Chang, Ricky Loynd, Robert West, Victor Dibia, Ahmed Awadallah, Ece Kamar, Rafah Hosn, Saleema Amershi. Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks. arXiv November 2024.
- Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, Ion Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks (arXiv June 2023).
- Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, Chenguang Zhu. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. EMNLP 2023 (arXiv March 2023).
- Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, Denny Zhou. Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024 (arXiv October 2023).
Engineering blogs and docs
- Anthropic (Erik Schluntz, Barry Zhang). Building effective agents. 19 December 2024.
- Anthropic (Jeremy Hadfield, Barry Zhang, Kenneth Lien, Florian Scholz, Jeremy Fox, Daniel Ford). How we built our multi-agent research system. Engineering blog, 13 June 2025.
- OpenAI. A practical guide to building agents. PDF, April 2025.
- LangChain. Graph API overview and Interrupts. LangGraph documentation, accessed October 2026.
- Temporal. Understanding Temporal. Temporal documentation, accessed October 2026.
- AWS (Aaron Sempf, Andrew Hooker). Agentic AI patterns and workflows on AWS: Workflow for routing. AWS Prescriptive Guidance, July 2025.
- Uber. uReview: Scalable, Trustworthy GenAI for Code Review at Uber. Engineering blog, 12 August 2025.
- Uber. QueryGPT: Natural Language to SQL Using Generative AI. Engineering blog, 19 September 2024.
- Uber. Enhanced Agentic-RAG: What If Chatbots Could Deliver Near-Human Precision?. Engineering blog, 29 May 2025.
- LinkedIn (Juan Pablo Bottaro, Karthik Ramgopal). Musings on building a Generative AI product. Engineering blog, 25 April 2024.
- LinkedIn (Albert Chen and colleagues). Practical text-to-SQL for data analytics. Engineering blog, 9 December 2024.
- DoorDash. Path to high-quality LLM-based Dasher support automation. Engineering blog, 17 September 2024; and Beyond Single Agents: How DoorDash is building a collaborative AI ecosystem. 11 November 2025.
- Spotify (Max Charas, Marc Bruggmann). Background Coding Agents: Predictable Results Through Strong Feedback Loops (Honk, Part 3). Engineering blog, 9 December 2025.
- Stripe (Alistair Gray). Minions: Stripe's one-shot, end-to-end coding agents. stripe.dev blog, 9 February 2026.
- Airbnb (Charles Covey-Brandt). Accelerating Large-Scale Test Migration with LLMs. Airbnb Tech Blog, March 2025.
- AWS (Aytul Arisoy Cholkar). Amazon Q Developer just reached a $260 million dollar milestone. AWS DevOps blog, 1 August 2024.
Code for this chapter
code/agents/ch4_chain.py: the three-step ticket chain under four conditions (strict or loose first step, with or without gates) against gpt-5-mini; results inresults/ch4_chain.json.code/agents/ch4_router.py,ch4_router_stats.pyandch4_grade.py: always-cheap, always-strong, cascade and classifier router on 24 questions, the cascade check read as a verifier, and the typed grader with its adversarial self-test.code/agents/ch4_parallel.py: plurality voting, best-of-n with a programmatic verifier, and sequential against concurrent wall-clock time.code/agents/ch4_evaluator.py: the evaluator-optimiser loop with a code checker or a model judge, judge-checker agreement, and a held-out fact check.code/agents/ch4_cost.py: chain reliability, routing cost, plurality voting with correlated errors, orchestrator cost, and the precision of a stop with an imperfect evaluator, each with its assumptions printed.code/agents/ch4_workflow.py: the complete composed workflow with typed state, schema-validated steps, budgets, checkpoints, resume, a human interrupt, an idempotent send and a trace;ch4_llm.pyis the shared API helper.code/agents/figs_ch4.pyandcode/agents/shots_ch4.py: the figures (with an automatic text-overlap check) and the paper excerpts.
Next
→ Chapter 5: Multi-agent systems
This chapter kept the control flow in code wherever it could and handed the model one decision at a time. Chapter 5 follows the orchestrator-workers pattern to its end: systems of several agents, each with its own loop, that hand work to one another, debate or vote. It asks the questions this chapter asked of every pattern (what changes, what it costs, how it fails and how you test it) of systems where the graph itself is partly written by the models, and it reads the evidence on when several agents beat one and when they only multiply the bill.








