Agentic Systems · Part 1 · Foundations

Chapter 1 · What an agent is, and when not to build one

What an agent is and when not to build one: the perceive-reason-act loop, levels of autonomy, reliability arithmetic, and what three companies learned.

Goal: by the end of this chapter you can say precisely what makes a system an agent rather than a model or a workflow, draw the loop that every agent runs, place any design on the spectrum from script to multi-agent system, do the arithmetic that tells you how many steps an agent can afford, and decide, for a given task, whether an agent is the right answer at all. You will also know where the idea came from and what three companies learned by running agents in production.


1.1 From one LLM call to a loop

You already know the basic move. You write a prompt, send it to a large language model, and get text back. One request, one response. The model does not remember the previous request unless you paste it in, it cannot look anything up, and it cannot do anything except produce text. This is a model call, and most useful LLM products are built from exactly this.

Now imagine a different kind of system. An alert fires at 14:06: the error rate of checkout-api has jumped to 12%. The system reads the alert. It decides that it needs more information, so it queries the log store for recent errors. The logs come back: 212 errors, almost all "connection pool exhausted", starting at 14:02, four minutes after a deploy of version 2.31. The system reads that, concludes the deploy is the likely cause, and opens a ticket proposing a rollback. Then it reports to the on-call engineer and stops.

Nothing in that story is a single call. The system took an observation (the alert), reasoned about it, chose an action (query the logs), observed the result, reasoned again, chose another action (open a ticket), and finally decided it was done. The same model was called three times, and each call saw everything that had happened so far. That repeating cycle is what makes the system an agent.

The idea is much older than language models. In the standard textbook on artificial intelligence, Artificial Intelligence: A Modern Approach, Stuart Russell and Peter Norvig define an agent as anything that perceives its environment through sensors and acts on it through actuators, and they describe a "rational agent" as one that chooses the action expected to do best given what it has perceived so far. A thermostat is an agent in that sense; so is a chess program; so is a robot. What is new since 2022 is that the "decide what to do next" part can be a general-purpose language model that reads the observations as text and writes its chosen action as text. That one substitution turned a forty-year-old concept into something you can build in an afternoon, and that is also why it is so easy to build badly.

contextwhat the model seesmodelreads, decides next steptool callthe actionenvironmentsearch_logsopen_ticketon-call engineerperceivereasonactobserve: the tool result is appended to the context, and the loop runs againfinal answer → user, stopwhen the model decides it is done (or the step budget runs out)ALERT: error rate 12%"check the logs first"search_logs(checkout-api)212 errors; deploy 13:58example, step 1 of the on-call assistant in §1.1: the grey line over each block is what that block held
The agent loop. The context (everything observed so far) is given to the model; the model picks a tool call; the tool acts on the environment; the result is appended to the context, and the loop runs again until the model gives a final answer or the step budget runs out. The grey text above each block shows step 1 of the on-call example.

Here is the loop in code. The whole thing is about fifty lines, and it runs with no API key because the "model" is a stub: a few if statements that return exactly the tool calls a real model would be expected to make for this one scenario. That is deliberate. The point of the script is the shape of the loop (where the model sits, what it sees, what the code does with its reply), not intelligence. Everything a real agent adds later (a real model, real tools, memory, tracing) plugs into one of these slots.

python
# code/agents/ch1_loop.py (the model is a stub; a real agent sends `messages` to an LLM here)
def model(messages):
    last = messages[-1]
    if last['role'] == 'user':
        return {'thought': 'An alert needs context first. Check recent errors for the service.',
                'tool': 'search_logs', 'args': {'service': 'checkout-api', 'minutes': 15}}
    if last['tool'] == 'search_logs':
        obs = last['content']
        return {'thought': f"{obs['errors']} errors, pool exhaustion, started 4 min after deploy ...",
                'tool': 'open_ticket', 'args': {'title': 'checkout-api: pool exhausted after v2.31', ...}}
    return {'thought': 'Ticket filed. Report back to the on-call engineer.', 'final': 'Opened OPS-4821 ...'}

def run_agent(task, max_steps=5):
    messages = [{'role': 'user', 'content': task}]
    for step in range(1, max_steps + 1):
        reply = model(messages)                                  # reason
        if 'final' in reply:                                     # the model decides it is done
            return reply['final']
        result = TOOLS[reply['tool']](**reply['args'])           # act
        messages.append({'role': 'tool', 'tool': reply['tool'], 'content': result})   # observe
    log('stopped: step budget exhausted')                        # a real loop always has a budget

Running it prints the trace: every thought, action and observation in order.

Terminal output of ch1_loop.py: the task, then three numbered steps, each with a THOUGHT line, an ACTION line calling search_logs or open_ticket, and an OBSERVE line with the tool result, ending with a FINAL line that reports the ticket

Three things in that output matter for the rest of the book. First, the model never touched the logs or the ticket system; it asked for things, and the loop did them. That is the boundary where permissions and guardrails live (Chapter 9). Second, every observation was appended to messages, so the second model call saw the alert and the log result. This growing list is the agent's context, and managing it is most of Chapter 2. Third, the loop had max_steps=5. A real model can loop forever on a task it cannot finish; a budget is not optional.

1.2 Model versus agent

Engineers often say "the agent hallucinated" or "the model did X" as if they were the same thing. They are not, and the difference decides where bugs live. A model is a function: text in, text out, stateless, with no goal and no ability to act. An agent is a system with a goal that runs over time; it has a loop, tools, some form of memory and the initiative to choose its next step. The model sits inside the agent as its decision-maker, the way a chess engine's evaluation function sits inside the engine.

A modelAn agent
Unit of workone calla task of many steps
Statenone between callscontext, memory, tool results carried forward
Can act on the worldno, it only writes textyes, through tools
Who chooses the next stepthe callerthe model, inside the loop
Stops whenthe response endsthe model says done, or a budget runs out
Failure looks likea wrong or made-up answera wrong action, a loop that never ends, a correct answer to the wrong sub-task
How you test itinput, expected outputthe whole trajectory: steps taken, tools used, final state
Costone prompt plus one completionthe sum over every turn, with a context that grows each turn

An analogy helps. A model is like a brilliant consultant on the phone who has amnesia between calls: ask a question, get a sharp answer, hang up, and the next call starts from nothing. An agent is the same consultant with a notebook, a phone directory and a to-do list, allowed to make calls, read the replies, write them down and keep going until the job is done. The notebook is memory. The directory is the tool set. The to-do list is the goal. The consultant's intelligence did not change; what changed is that it can now do things and can now do them wrong in ways that persist.

A model: one callAn agent: a goal pursued over many stepspromptmodeltextprompt in, text outno memory of earlier callsno actions, no goal of its owngoal: investigate the alert, propose a fixmodelreasonstoolsact, observememorystate so faractobserveread, writethe model is one part;the loop, tools andmemory are the restmany steps, state carried between them,and the model picks each next step
Left: a model computes one output from one prompt. Right: an agent wraps the same model in a loop with a goal, a tool set and memory, and pursues the goal over many steps; the model is one component, not the whole.

The practical consequence: when an agent fails, ask where in the loop it failed. Did the model misread an observation (a model problem)? Was the right tool missing or badly described (a tool problem)? Did the loop keep going after the task was done (a loop problem)? Was the needed fact pushed out of a too-long context (a memory problem)? The same symptom, a wrong final answer, can come from any of them, and only the trace tells you which.

1.3 Workflows versus agents

Here is the most useful distinction in the field, and the one most often blurred in product pitches. Many systems call a language model several times, pass data between the calls and use tools, yet are not agents, because the code decides the sequence. Anthropic's December 2024 post "Building effective agents" draws the line clearly.

The same post makes a second point that this whole chapter is built on: start simple.

OpenAI's practical guide to building agents (a PDF published in April 2025) reaches the same place from the other side. It defines agents as systems that "independently accomplish tasks on your behalf", says plainly that applications which use an LLM but do not let it control the workflow ("simple chatbots, single-turn LLMs, or sentiment classifiers") are not agents, and then gives three criteria for when an agent is worth it.

Put the two posts together and you get a spectrum rather than a binary.

scriptcode decidesevery stepcron job, a rules engineworkflow + LLM stepscode decides the path;the LLM fills in stepssummarise, then translateagentthe LLM decideseach next stepand when to stopon-call assistantmulti-agentseveral LLMs decide;one coordinatesresearch systemmore autonomy, more cost, more variance between runs, harder to test →who decides the next step:codecode, with LLM helpthe LLM(s)
The spectrum from a plain script to a multi-agent system. Moving right, the model takes over more of the decision about which step comes next; cost, variance between runs and testing effort all rise with it.
  • Script. Code decides everything. A cron job, a rules engine, a SQL report. Deterministic, cheap, testable, and the right answer far more often than engineers who have just discovered agents want to admit.
  • Workflow with LLM steps. Code decides the path; the model fills in steps that need language: classify this email, summarise this document, extract these fields. Most production "AI features" live here, and they should.
  • Agent. The model decides the next step, including when to stop. Needed when the number and order of steps depends on what is discovered along the way: debugging, research, open-ended support, operating a browser.
  • Multi-agent system. Several agents, usually one coordinating the others. Needed rarely, and mostly when the work parallelises or exceeds one context window. Chapter 5 is about when this helps and the many ways it fails.
Workflow: the path is written in codeAgent: the model chooses the pathinput: a support emailLLM: classify the topiccode: look up the accountLLM: draft a replycode: queue for reviewsame five steps on every run; easy to test step by stepmodelpicks a toolread_email1lookup_account2search_kb3draft_reply4stop: answerorder, count and stop chosen at run time; test the whole trajectory
Left: a workflow runs the same code path on every input and calls the model inside fixed steps, so each step can be tested on its own. Right: an agent is given a set of tools and chooses the order, the number of steps and the stopping point at run time, so the whole trajectory must be tested.

The figure's right side explains why the rest of this book spends four chapters on evaluation and observability. In the workflow, you can unit-test "classify the topic" with a hundred labelled emails. In the agent, there is no fixed step to test: on one input it reads the email then looks up the account; on another it searches the knowledge base first; on a third it loops twice. You can only test the behaviour, over many runs, and you can only debug it from the trace.

1.4 Where the idea came from

Five papers and two engineering posts, from October 2022 to June 2025, cover the whole arc from "a model can reason and act in turns" to "here is how multi-agent systems fail in production". Later chapters read several of them closely; this section gives one paragraph each so you know the names.

202320242025ReActOct 2022reason, act, observeToolformerFeb 2023a model taught to call APIsReflexionMar 2023verbal self-reflection, retryGenerative AgentsApr 202325 agents in a small townBuilding effective agentsDec 2024workflows versus agentsMASTMar 2025why multi-agent systems failMulti-agent research systemJun 2025orchestrator and subagentsresearch paperengineering blog post
Timeline. ReAct (October 2022) made the reason-act-observe loop explicit; Toolformer, Reflexion and Generative Agents followed within six months in 2023; the production write-ups, workflows versus agents, the failure taxonomy and the multi-agent research system, arrived in 2024 and 2025.

ReAct (October 2022). Shunyu Yao and colleagues at Princeton and Google Research noticed that two lines of work had been kept apart: prompting a model to reason step by step (chain of thought), and prompting it to emit actions. ReAct interleaves them. The model writes a thought, then an action, then reads the observation, then writes the next thought. This is the loop of Section 1.1 written down as a prompting format, and it is the direct ancestor of every tool-using agent since. Chapter 3 reads the paper in full.

Toolformer (February 2023). Timo Schick and colleagues at Meta AI asked a different question: instead of prompting a model to use tools, could a model learn to call them? Toolformer takes a plain language model and teaches it, from a handful of examples per tool and a self-supervised filtering trick, to insert API calls (a calculator, a search engine, a calendar) into its own text where they reduce its prediction error. It matters here because it shows that tool use is not a prompt hack: it can be a trained ability, and the "function calling" features now built into commercial models are the productised descendants of this idea.

Reflexion (March 2023). Noah Shinn and colleagues took the loop one level up. When an agent fails a task, Reflexion asks the model to write a short verbal reflection on why it failed, stores that text in an episodic memory, and includes it in the next attempt. No weights are updated; the "learning" is in the text. The result was large gains on coding and reasoning benchmarks, and the pattern (try, critique, retry with the critique in context) is now a standard tool in the box. Chapter 3 covers it as the reflection pattern.

Generative Agents (April 2023). Joon Sung Park and colleagues at Stanford and Google put twenty-five agents in a small simulated town, each with a memory stream of everything it had observed, a mechanism to retrieve relevant memories, and a step that periodically turned memories into higher-level reflections and plans. The agents woke up, cooked breakfast, went to work, held conversations, and, when one was told she wanted to throw a party, spread the invitation through the town so that others showed up. It is a research demonstration rather than a production system, but it established the memory architecture (observe, store, retrieve, reflect, plan) that Chapter 2 builds on, and it is the paper that made "agents" a mainstream word.

"Why Do Multi-Agent LLM Systems Fail?" (March 2025). Two years after the burst of 2023, Mert Cemri, Melissa Pan, Shuyi Yang and colleagues at UC Berkeley asked the uncomfortable question in their title. They collected execution traces from popular open-source multi-agent frameworks, had experts annotate where each failed, and built a taxonomy (MAST) of fourteen failure modes in three groups. The version of the paper published at NeurIPS 2025 reports over 1,600 annotated traces across seven frameworks, failure rates between 41% and 86.7% on the systems studied, and a split of failures into system design issues (44.2%), misalignment between agents (32.3%) and weak task verification (23.5%).

Two engineering posts close the timeline, and Section 1.6 reads both: Anthropic's "Building effective agents" (December 2024), which gave the field the workflow-versus-agent vocabulary, and Anthropic's account of its multi-agent research system (June 2025), which is the most detailed public description of a production orchestrator-and-subagents design, including its evaluation.

1.5 Levels of autonomy and the economics of agents

Two questions decide whether an agent design is viable before a single line is written: how much are you letting it do without asking, and what does each step cost in reliability, money and time.

The autonomy ladder

LevelThe agent...A person...ExampleWhat a wrong step costs
1. Suggestreads, reasons, proposesreads the proposal and actsdrafts a reply to a support emaila bad draft, caught before sending
2. Act with approvalprepares an action and asksclicks approve or rejectproposes a rollback and waitsone click of attention per action
3. Act and reportacts, then says what it didreviews the report, can undorolls back, then posts to the channela wrong action, visible quickly, reversible if designed so
4. Fully autonomousacts; nobody looksis alerted only on failuretriages and closes tickets overnightwrong actions that compound before anyone sees them
1 suggestdrafts a reply;a person sends ita wrong draft is read,not acted on2 act with approvalproposes a rollback;a person clicks yesa person is thelast check3 act and reportrolls back, thenposts what it didwrong actions happenand are seen fast4 fully autonomousacts; nobody looksunless alertedwrong actionscompound unseenmore autonomy: less human time per task, bigger blast radius when a step goes wrong →
Four levels of autonomy, from suggesting to acting without review. Each level up saves human time per task and widens the blast radius of a single wrong step.

The ladder is not a maturity model where level 4 is the goal. Most production agents today, including the three in Section 1.6, run at levels 1 and 2, and that is a design decision, not a shortcoming. The right level for an action depends on two things: how reliable the agent is at that action (measured, not assumed) and how reversible the action is. Reading logs can be level 4 from day one. Deleting records should be level 2 until you have the evidence to argue otherwise.

The reliability arithmetic

An agent that takes ten steps must get all ten right. If each step succeeds independently with probability pp, the whole task succeeds with probability

P(task)=pnP(\text{task}) = p^{n}

for nn steps. Independence is an approximation (errors often correlate, and a good agent can recover from some), but the shape of the result is what matters, and it is brutal. The script ch1_math.py computes the table.

Terminal output of ch1_math.py: a table of success probability for p of 0.9, 0.95, 0.99 and 0.999 across 1, 5, 10, 20 and 50 steps; the number of steps each p affords before success falls below 90 percent and 50 percent; a token-cost table for one call, an eight-step agent and a lead agent with four subagents at an example price; and a latency line

step reliability pp5 steps10 steps20 steps50 stepssteps that keep success above 90%
0.9059.0%34.9%12.2%0.5%1
0.9577.4%59.9%35.8%7.7%2
0.9995.1%90.4%81.8%60.5%10
0.99999.5%99.0%98.0%95.1%105
0%25%50%75%100%90% target151020304050p = 0.9, 1 steps: 90.0%p = 0.9, 5 steps: 59.0%p = 0.9, 10 steps: 34.9%p = 0.9, 20 steps: 12.2%p = 0.9, 50 steps: 0.5%p = 0.95, 1 steps: 95.0%p = 0.95, 5 steps: 77.4%p = 0.95, 10 steps: 59.9%p = 0.95, 20 steps: 35.8%p = 0.95, 50 steps: 7.7%p = 0.99, 1 steps: 99.0%p = 0.99, 5 steps: 95.1%p = 0.99, 10 steps: 90.4%p = 0.99, 20 steps: 81.8%p = 0.99, 50 steps: 60.5%p = 0.999, 1 steps: 99.9%p = 0.999, 5 steps: 99.5%p = 0.999, 10 steps: 99.0%p = 0.999, 20 steps: 98.0%p = 0.999, 50 steps: 95.1%p = 0.999: 95.1% at 50 stepsp = 0.99: 60.5% at 50 stepsp = 0.95: 7.7% at 50 stepsp = 0.9: 0.5% at 50 stepsnumber of steps n that must all succeedP(task succeeds) = pⁿ
Probability that a task of n steps succeeds when each step succeeds with probability p. At p equal to 0.90, a ten-step task succeeds about a third of the time; a step reliability of 0.99 buys ten steps above the 90 percent line; only 0.999 stays above it for a hundred steps.

Read the table as a budget. A model that is right 90% of the time on each step, which sounds respectable, gives you a ten-step agent that fails two times in three. To run ten steps with 90% task success you need every step at 99%. To run fifty steps you need 99.9%, which is the territory of carefully constrained tools, strict output formats and retries, not of free-form reasoning. This is why so many agent demos work beautifully and so few agent products ship: the demo is one run, and the product is pnp^{n} over ten thousand runs.

Three design moves follow directly from the formula, and each is a later chapter:

  1. Reduce nn. Fewer, bigger, better-designed tools (Chapter 2); fixed workflow steps for the parts that do not need judgement (Chapter 4).
  2. Raise pp. Clear instructions, good tool descriptions, structured outputs, and the reasoning patterns of Chapter 3; evals that tell you which kind of step is weak (Chapter 6).
  3. Break the independence. Verification steps, retries and self-checks turn a chain of must-succeed steps into one with recovery paths, so a single slip no longer sinks the run (Chapters 3 and 9).

Token and latency cost

Every turn of the loop sends the whole context to the model again. An eight-step agent whose context averages 3,000 input tokens per step therefore sends about 24,000 input tokens, against 3,000 for a single call. To make that concrete, ch1_math.py prices it at an example price of $3 per million input tokens and $15 per million output tokens. These are round numbers chosen for the arithmetic, not a quote from any provider; substitute your own.

systemmodel callsinput tokensoutput tokenscost per taskcost per 10,000 tasks
one LLM call13,000500$0.0165$165
agent, 8 steps824,0003,200$0.1200$1,200
lead agent plus 4 subagents of 8 steps33120,00016,000$0.6000$6,000
Cost of 10,000 tasks at the EXAMPLE price of $3 / M input and $15 / M output tokensone LLM callone LLM call: $165$165agent, 8 stepsagent, 8 steps: $1,200$1,200lead + 4 subagentslead + 4 subagents: $6,000$6,000model callsinput tokenslatency at 2 s per callone LLM call13,0002 sagent, 8 steps824,00016 slead + 4 subagents33120,00032 s (subagents in parallel)
The same task as one model call, as an eight-step agent, and as a lead agent with four subagents of eight steps each, at an example price. The agent costs about seven times the single call and the multi-agent system about thirty-six times; the sequential agent also takes about eight times as long.

The agent costs 7.3 times the single call and the multi-agent system 36 times, before counting retries. Latency stacks the same way: if a model call takes about two seconds, eight sequential steps take sixteen, plus tool time. Anthropic's production numbers, in Section 1.6, are in the same range: agents used about four times the tokens of a chat interaction and multi-agent systems about fifteen times.

None of this says agents are too expensive. It says the task has to be worth it: an agent that saves an engineer twenty minutes of log-reading is worth far more than $0.12. But it does mean that every design decision about steps, context length and model size is a cost decision, and that you need to see those numbers per run, which is what tracing (Chapter 8) gives you.

This arithmetic is why evals and tracing sit at the heart of this book rather than in an appendix. pp cannot be guessed; it must be measured per kind of step, which is an eval. nn and the token count cannot be reasoned about from the code; they emerge at run time, which is a trace. An agent without both is a system whose reliability and cost are unknown by construction.

1.6 What companies actually run

Three systems, chosen because their teams published enough detail to learn from and because the claims can be checked against the published text. For each: what the agent does, the architecture in one sentence, how it is evaluated, and what the team said went wrong or surprised them. Where a post gives no number, this section says so rather than inventing one.

(a) Uber's Genie: an on-call copilot in Slack

Uber's engineering teams run Slack channels where internal users ask for help with platforms such as Michelangelo, Uber's machine-learning platform. The October 2024 post "Genie: Uber's Gen AI On-Call Copilot" describes the bot built to answer those questions.

What it does. A user posts a question in a Slack channel; Genie answers in the thread with a response grounded in Uber's internal documentation, with citations, and offers feedback buttons. If the answer does not help, the on-call engineer picks it up as before.

Architecture in one sentence. Retrieval-augmented generation wrapped in a Slack bot: documents are chunked and embedded into a vector database; a back-end "Knowledge Service" embeds the incoming question, fetches the most relevant chunks, and sends them with the question to an LLM through Uber's Michelangelo gateway. The post says the team chose RAG over fine-tuning because it needs no training examples to start, which "reduced the time to market".

Genie (Uber, 2024), as the engineering blog describes itinternal docsthe knowledge sourceschunk + embedofflinevector databasechunks by embeddingSlack questiona user in asupport channelKnowledge Serviceembed the question,fetch relevant chunksLLMvia the Michelangelogateway; prompt = chunksanswer in Slackwith citationsqueryuser feedbackResolved / Helpful /Not Helpful / Not RelevantevaluationLLM as a judge on custom metrics (hallucination, relevancy);channel owners use the results to improve their docsbetter docs
Genie as the Uber blog describes it. Internal documents are chunked and embedded into a vector database; a Slack question is embedded, matched to chunks and answered by an LLM with citations; user feedback and LLM-as-a-judge evaluations feed back to the owners of the documents.

Be honest about the label. As described in the 2024 post, Genie is closer to a workflow than to an agent in the Section 1.3 sense: the path (embed, retrieve, generate, reply) is fixed in code, and the model fills in the generation step. It is in this chapter because its evaluation loop is exactly what agentic systems need, and because it shows the simplest-solution principle in practice: a fixed pipeline answered tens of thousands of questions before anyone needed to let the model choose steps.

How they evaluate it. Two mechanisms. First, feedback buttons on every answer.

Second, channel owners can run custom offline evaluations: the post describes an evaluator that fetches a specified prompt, runs "LLM as a Judge", and extracts the metrics the owner cares about, such as hallucination and answer relevancy, so that teams can improve the documents the bot reads.

The numbers the post reports.

What they learned. The post does not give a failure analysis, so this section will not invent one. What it does show is the design pattern worth copying: launch a fixed pipeline, instrument every answer with a cheap label, give the owners of the knowledge a way to measure their own slice, and let the measured helpfulness rate drive what to improve.

(b) Anthropic's multi-agent research system

In June 2025 Anthropic published a detailed account of how its Research feature works: a system that, given an open question, searches the web and internal sources for several minutes and returns a cited report.

What it does. Takes a research question, plans an approach, runs several searches in parallel, reads the results, decides whether more is needed, and writes a report with citations.

Architecture in one sentence. An orchestrator-worker design: a lead agent (Claude Opus 4) plans the research, saves its plan to memory, spawns subagents (Claude Sonnet 4) that each search a sub-question in parallel with their own tools and context, synthesises what they return, and hands the draft to a citation agent that attaches sources.

Anthropic's research system (2025): an orchestrator and parallel subagentsuser querya research questionlead agent (Opus 4)plans, splits the question intosubtasks, later synthesisesmemorythe saved plansubagent 1 (Sonnet 4)searches, reads, reportsweb searchsubagent 2 (Sonnet 4)searches, reads, reportsweb searchsubagent 3 (Sonnet 4)searches, reads, reportsweb searchgrey: subtasks out,in parallel;blue: findings backcitation agent → reportafter synthesis
The orchestrator-worker design of the multi-agent research system as the post describes it: a lead agent plans and delegates, several subagents search in parallel with their own tools, the lead synthesises their findings, and a citation agent attaches sources.

The post is unusually frank about cost.

How they evaluate it. The post's section on evaluation is the most useful part for this book, and Chapter 7 returns to it.

The same section describes the rest of the method: an LLM judge scoring each answer on a rubric (factual accuracy, citation accuracy, completeness, source quality, tool efficiency), where "a single LLM call with a single prompt outputting scores from 0.0-1.0 and a pass-fail grade was the most consistent"; and human testers, because "people testing agents find edge cases that evals miss", such as the agents' early preference for SEO-optimised content farms over authoritative sources.

What they learned. The post lists production problems that any long-running agent will meet: agents are stateful and errors compound, so a minor failure can derail a run; the team added the ability to resume from where an error occurred rather than restart; it used "rainbow deployments" that keep old and new versions running side by side so that an in-flight agent is not disrupted by a deploy; and it found that full tracing of agent decision patterns was needed to diagnose failures without reading users' conversations. Every one of those is a chapter of this book.

(c) LinkedIn's generative AI product

The third system is older and less agentic than the second, and is here for its honesty about evaluation. In April 2024 two LinkedIn engineers, Juan Pablo Bottaro and Karthik R., published "Musings on Building a Generative AI Product", describing the feature that lets a member ask questions about a job post or a feed item and get a tailored answer.

What it does. A member asks, for example, how well they fit a job; the system routes the question, gathers the member's profile and the job through internal APIs, and writes an answer.

Architecture in one sentence. A router decides whether the question is in scope and which topic-specific "AI agent" should handle it; that agent runs a retrieval step (recall-oriented: internal APIs and search) and a generation step (precision-oriented: write from what was found), with the internal APIs wrapped as "skills" the model can call.

LinkedIn's generative AI product (2024): route, retrieve, generate, then measuremember questionin the feed or on a jobroutingin scope? which agent?job assessment agentcompany understanding agenttakeaways for posts agentone such path per agentretrievalinternal APIs andsearch, recall firstgenerationprecision: writefrom what was foundquality measurementlinguists score up to 500 conversations a day:overall quality, hallucination rate,Responsible AI violations, coherence, stylethe blog: routing and retrieval were tuned on dev sets, like classifiers;generation reached 80% fast, and the last 20% took most of the work
LinkedIn's design as the blog describes it: a router decides whether a question is in scope and which topic agent handles it; that agent retrieves from internal APIs and search, then generates an answer; a linguist team scores hundreds of conversations a day.

In Section 1.3 terms this is a routing workflow whose branches call tools; the authors use the word "agent" for each branch. The pattern (classify, dispatch, retrieve, generate) is Chapter 4's routing pattern, and it is the most common production architecture in the industry.

How they evaluate it, and what was hard.

What went wrong. The post describes the model producing malformed structured output when calling the internal APIs: about 10% of responses had parameter mistakes, which the team reduced to about 0.01% by analysing the common errors, writing code to detect and patch them before parsing, and adding hints to the prompt.

The post also reports a latency trade-off that every agent designer meets: chain-of-thought reasoning improved quality and reduced hallucination but added tokens "that the member never sees", increasing perceived latency; and that watching only time-to-first-token hid a degradation in time between tokens that set off alerts during a public ramp and forced a capacity increase. Measure both.

What the three have in common

Uber GenieAnthropic ResearchLinkedIn
Spectrum position (§1.3)workflow (RAG pipeline)multi-agentrouting workflow with tool-calling branches
Autonomy level (§1.5)1: suggests an answer, human fallback1: produces a report1: produces an answer
Online evalfour-label feedback buttonsuser feedback (described, no numbers given)not described
Offline evalLLM-as-a-judge on owner-chosen metrics~20 queries to start, LLM judge with a rubric, human testerslinguist team, up to 500 conversations a day, five metrics
Published numbers45k questions/month, 154 channels, 70k answered, 48.9% helpful, 13k hours+90.2% versus single agent (internal), 4× and 15× tokens~10% to ~0.01% payload errors; "80% fast, last 20% most of the work"
Hardest part, in their words(not stated)evaluation, compounding errors, deployment of stateful agentsevaluation guidelines and annotation; the last 20% of quality

Notice what is absent: none of the three runs at autonomy level 3 or 4, none of the three skipped evaluation, and two of the three are workflows rather than agents by the strict definition. The companies with the most experience chose the simplest design that worked and put their effort into measurement.

1.7 When NOT to build an agent

Here is the checklist this chapter has been building towards. Go through it before writing a prompt. An agent is the wrong answer if any of the following holds.

  1. Deterministic logic is enough. If you can write the steps as code and they are the same for every input, write the code. A model adds cost, latency and variance, and removes testability, for no gain.
  2. The cost or latency budget does not fit nn steps. If the task must complete in two seconds or cost under a cent, an eight-step agent is out before you start. Do the Section 1.5 arithmetic with your own numbers.
  3. The accuracy requirement is above what pnp^{n} allows. If the task must succeed 99% of the time and you have ten steps, you need p≈0.999p \approx 0.999 per step. If you cannot measure pp or cannot reach it, reduce nn or use a workflow.
  4. There is no way to evaluate it. If you cannot say what a correct outcome looks like for a few dozen representative inputs, you cannot tell whether the agent works, cannot tell whether a change helped, and cannot tell when it has silently got worse. Build the eval set first; it will also tell you whether you needed an agent.
  5. The actions are irreversible and there is no approval step. Sending money, deleting data, emailing customers, changing production configuration. Either design the approval step (level 2) and the undo path, or do not let the model act.
Can plain code do it reliably?yesWrite the code. No model needed.noAre the steps the same on every run?yesA workflow: a fixed path, the model inside fuzzy stepsnoCan you measure success (an eval set)?noNot yet. Collect examples and build the eval first.yesDoes the step budget fit cost and latency?noFewer steps, a cheaper model, or a workflow.yesWould a wrong action be hard to undo?yesAn agent that acts only with approval (level 2)noAn agent that acts and reports (level 3)raise autonomy only as the evals and traces earn it
A decision flow for whether to build an agent. Most tasks leave at the first two questions; an agent is the answer only when the steps vary between runs, success can be measured, the budget fits the step count, and wrong actions can be undone or approved.

Three worked scenarios, deliberately mundane:

Scenario 1: turn the office lights off at 7 pm. Checklist item 1 fails immediately: a cron job and one API call do this perfectly, for free, forever. There is no judgement, no variation between runs and no natural language. Anyone proposing an agent here has confused "uses AI" with "is good". Answer: a script.

Scenario 2: triage incoming customer emails. Item 1 passes: emails are unstructured, and the decision (billing? bug? refund? spam?) needs language understanding. Item 2 passes: a few seconds and a fraction of a cent per email is fine. Item 3 and 4 need work: you need a few hundred labelled emails to measure the classifier's accuracy, and you need to decide what accuracy is acceptable given that a human reads the result. Item 5 passes if the action is "put in a queue with a suggested label" (level 1) rather than "reply automatically". But look at the path: classify, look up the account, draft a reply, queue for review. It is the same every time. Answer: a routing workflow with model steps, at autonomy level 1. Not an agent. Promote the "draft a reply" step to a small agent with tools (search the knowledge base, read the order history) only when the eval shows the fixed path is what limits quality.

Scenario 3: flag transactions over a threshold. Item 1 fails: amount > threshold is one line of SQL, and a model would be slower, dearer and occasionally wrong at comparing two numbers. Answer: a rule. The interesting version is the one OpenAI's guide uses: fraud investigation, where the rule has flagged a transaction and someone must now read the account history, look for patterns and decide. That has variable steps, unstructured evidence and judgement, and could justify an agent at level 1 (write up the case for a human analyst) once there is an eval set of past investigations to measure it against.

1.8 The map of this book

The chapters that follow are organised as the layers of a production agent, from the inside out.

Ch 1 and 2the loop, tools,memory, contextCh 3reasoning patternsCh 4workflow patternsCh 5multi-agentCh 6 and 7evalsCh 8tracingCh 9guardrailsthe outer ring is most of this bookCh 10: capstoneone agent throughevery layercentre: what an agent is made ofmiddle: how its steps are arrangedouter: how it is measured, watchedand kept safeThe ten chapters as layers
The map of the book. The loop, tools, memory and context sit at the centre (Chapters 1 and 2); the reasoning, workflow and multi-agent patterns are arranged around them (3 to 5); evals, tracing and guardrails form the outer ring (6 to 9), which is most of the book; the capstone (10) builds one agent through every layer.

At the centre is what an agent is made of: this chapter's loop, and Chapter 2's building blocks (model, tools, instructions, memory, state) with a close look at how to design a tool a model can actually use and how to manage what is in the context. Around that are the patterns: Chapter 3 on how a single agent reasons (ReAct, plan-and-execute, reflection, tree search, and what the papers measured), Chapter 4 on the workflow patterns that fix a path in code (chaining, routing, parallel fan-out, orchestrator-workers, evaluator-optimiser), and Chapter 5 on multi-agent systems and, with the MAST taxonomy in hand, the ways they fail.

The outer ring is the largest, because it is where production systems live or die: Chapters 6 and 7 on evaluation (golden sets, error analysis, LLM-as-a-judge and its calibration, trajectory evals, public benchmarks, and then evals in production: offline suites, online checks, regression gates in CI, and how the companies in this chapter did it), Chapter 8 on observability (traces and spans for agents, what to log, replaying a run, cost and latency dashboards) and Chapter 9 on guardrails and human oversight (input and output checks, tool permissions, sandboxing, prompt injection, approval steps, kill switches). Chapter 10 builds one small production-shaped agent through every layer, with the code in the repository, and ends with a cheatsheet and a glossary.

Exercises

  1. Take a system you have built or used that is described as "AI-powered". Draw its flowchart. Is every box and arrow known before a run starts? Place it on the spectrum of Section 1.3 and give the autonomy level of each action it takes.
  2. For the on-call assistant of Section 1.1, list five distinct ways a single turn of the loop could fail (a wrong tool, bad arguments, misread observation, and so on). For each, say whether it is a model, tool, loop or memory problem, and what line in the trace would reveal it.
  3. Your agent needs 12 steps and must succeed on 95% of tasks. What per-step reliability do you need? Which of the three design moves in Section 1.5 would you try first to get there, and why?
  4. Re-price the cost table in Section 1.5 with your own provider's prices and a context that grows by 1,500 tokens per step rather than staying at 3,000. At what step count does the agent cost more than a human doing the task for ten minutes at your organisation's loaded hourly rate?
  5. Pick one of the three company systems in Section 1.6 and write the eval set you would build for it: ten representative inputs, the definition of success for each, and who or what would judge it. Where would an LLM judge be acceptable and where would you insist on a human?

Key takeaways

  • A model computes one output from one prompt and is stateless. An agent is a loop around a model: observe, reason, act, observe again, until the model says it is done or a budget runs out. The agent is defined by the loop, not the model.
  • The loop has four slots, context, model, tool call and environment, and every agent failure lives in one of them. Only the trace tells you which.
  • Workflows fix the path in code and call the model inside steps; agents let the model choose the path at run time. The test is "who decides the next step", not "does it use tools". Most production AI features are, and should be, workflows.
  • Both Anthropic and OpenAI say the same thing: find the simplest solution, and build an agent only where the path cannot be written down in advance (variable steps, judgement, unstructured data).
  • The idea runs from ReAct (interleaved reasoning and acting, 2022) through Toolformer (learned tool use), Reflexion (verbal self-reflection in memory) and Generative Agents (a memory architecture for long-running behaviour) to the 2025 failure taxonomy, where most multi-agent failures were system design problems, not model problems.
  • Autonomy is a ladder (suggest, act with approval, act and report, fully autonomous), and the right rung depends on measured reliability and reversibility. The three production systems in this chapter all run at level 1.
  • Reliability compounds: a task of nn steps at per-step reliability pp succeeds with probability pnp^{n}. At p=0.9p = 0.9 a ten-step task fails two times in three; ten steps above 90% needs p=0.99p = 0.99 on every step.
  • Cost compounds too: an eight-step agent costs about seven times a single call in the example and takes about eight times as long; Anthropic measured 4× tokens for agents and 15× for multi-agent systems in production. The task has to be worth it.
  • The companies that run agents put their effort into evaluation: four-label feedback buttons, twenty real queries and an LLM judge to start, five hundred human-scored conversations a day. "80% was fast; the last 20% took most of the work."
  • Do not build an agent when code can do it, when the budget does not fit the step count, when the accuracy requirement is above what pnp^{n} allows, when you cannot evaluate it, or when the actions are irreversible with no approval step. Build the workflow first; it gives you the baseline and the eval set.

References

Papers and books

  1. Stuart Russell and Peter Norvig. Artificial Intelligence: A Modern Approach, 4th edition. Pearson 2020. (Chapter 2, "Intelligent Agents".)
  2. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, Yuan Cao. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023 (arXiv October 2022).
  3. Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, Thomas Scialom. Toolformer: Language Models Can Teach Themselves to Use Tools. NeurIPS 2023 (arXiv February 2023).
  4. Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao. Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023 (arXiv March 2023).
  5. Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein. Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023 (arXiv April 2023).
  6. Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica. Why Do Multi-Agent LLM Systems Fail?. NeurIPS 2025 Datasets and Benchmarks track (arXiv March 2025).

Engineering blogs and docs

  1. Anthropic. Building effective agents. Research blog, 19 December 2024.
  2. OpenAI. A practical guide to building agents. PDF guide, April 2025.
  3. Nicholas Marcott, Eduards Sidorovics, Paarth Chothani, Xiyuan Feng, Chun Zhu, Meghana Somasundara, Kailiang Fu, Jonathan Li. Genie: Uber's Gen AI On-Call Copilot. Uber Engineering blog, 10 October 2024.
  4. Anthropic. How we built our multi-agent research system. Engineering blog, 13 June 2025.
  5. Juan Pablo Bottaro and Karthik R.. Musings on Building a Generative AI Product. LinkedIn Engineering blog, 25 April 2024.

Code for this chapter

  1. code/agents/ch1_loop.py: the fifty-line agent loop with a stubbed model and two fake tools; its output is results/ch1_loop_stdout.txt.
  2. code/agents/ch1_math.py: the reliability table (pnp^{n}), the step budgets, and the token-cost comparison at an example price; results in results/ch1_math.json.
  3. code/agents/figs_ch1.py and code/agents/shots_ch1.py: the figures and the paper excerpts.

Next

→ Chapter 2: The building blocks

Chapter 2 opens the agent up. It takes the four slots of the loop (model, tools, instructions, memory and state) one at a time, shows how to design a tool so that a model calls it correctly (the "which, when, what arguments, what to do with the result" of Toolformer), and introduces context engineering: deciding what goes into the model's input on each turn, what is summarised, what is stored outside and retrieved, and what is thrown away. It is the chapter that turns the stubbed model() of this chapter into something real.