Agentic Systems · Part 1 · Foundations

Chapter 2 · The building blocks: model, tools, instructions, memory and state

The five building blocks of an agent: model, tools, instructions, memory and state, how to design tools a model can call, and context engineering.

Goal: by the end of this chapter you can take the loop of Chapter 1 apart into its five blocks, say what each one is responsible for and what breaks when it is weak, choose a model for an agent with a measurement rather than a feeling, design a tool a model will actually call correctly, write instructions that behave like architecture rather than like a wish list, give an agent the right kind of memory for the right length of time, keep its state so that a run can be resumed, and decide, token by token, what belongs in the context window. Every choice comes with its trade-offs, a way to test whether it was right, and what tends to go wrong in production.


2.1 The five blocks and how they connect

Chapter 1 drew the agent as a loop: context in, decision out, tool runs, observation back, repeat. That picture is right but it hides the engineering. When you build a real agent you are making five separate kinds of decision, and when one fails in production you need to know which of the five it was. This chapter calls them the building blocks.

The five blocks of an agent and what flows between theminstructionsshape: role, rules, formatmemoryremember: facts, notes, storemodelreason: read the context,pick the next actiontoolsact: run the call, returnan observationstate (the loop)control: step, plan, pendingcall, budget, errorsevery turnretrievewrite notes, factstool callobservationnext turn? budget left?errors, retriesenvironmentcalendar, mail, CRM, fileswhat each block is for: the model decides, the tools do, the instructions constrain,the memory carries knowledge forward, the state says where in the task the loop is
The five building blocks and what flows between them. Instructions and retrieved memory are assembled, with the history, into the context the model reads on every turn; the model emits a tool call; the tool acts on the environment and returns an observation; the loop keeps the state (which step, what plan, which call is pending, how much budget is left, which errors have happened) and decides whether to call the model again.

Each block owns one question.

BlockThe question it answersMade ofTypical failure when it is weak or missing
Model"Given all this, what should happen next?"an LLM behind an API or on your own hardwarewrong tool, wrong arguments, misread observation, gives up or loops
Tools"How does the decision become an effect in the world?"functions with a name, a description and a schemathe model cannot act, acts on the wrong thing, or drowns in the result
Instructions"What is this agent for, and what must it never do?"the system prompt: role, rules, formats, examples, stop conditionsdrifts off task, ignores constraints, output that no parser accepts
Memory"What does the agent know from before this turn, this run, or earlier runs?"the context window, a scratchpad, a store it reads and writesrepeats work, forgets the user's constraint, contradicts itself
State"Where in the task is the loop, and can it pick up where it left off?"a small object: step, plan, pending call, budget, errors, statuscannot resume after a crash, retries blindly, never knows it is done

Two of these are easy to confuse, and the confusion causes real bugs. Memory is knowledge: what the agent has learned or been told. State is control: where the loop is in the task. "The user prefers mornings" is memory. "We are at step 4, the email tool just failed once, and there are four steps of budget left" is state. A system that keeps only memory cannot resume a crashed run safely; a system that keeps only state cannot remember what it was told. Section 2.6 draws the line in detail.

The blocks are connected by one artefact: the context the model receives on each turn. Instructions go in first and stay fixed. Tool definitions go in so the model knows what it may call. Memory contributes whatever was retrieved. The history of the run so far, including every observation, is appended. The last observation is the newest thing in it. The model reads all of that and writes one decision. This is why Section 2.7 on context engineering sits at the end of the chapter: it is where the other four blocks meet, and it is the budget they all draw on.

Chapter 1 gave the arithmetic of why reliability compounds across steps. This chapter is about the levers that raise per-step reliability, and almost all of them live in these blocks: a model that calls tools correctly, tools that are hard to misuse, instructions that remove ambiguity, memory that puts the right fact in front of the model, and state that makes a failed step recoverable rather than fatal.

Discussion

A team designing an agent should argue about these before writing a line:

  1. Which block are we weakest at, and how would we know? Most teams assume the model is the weak block and reach for a bigger one. In the production write-ups of Chapter 1, the weak blocks were tools (LinkedIn's malformed payloads) and instructions and verification (the MAST taxonomy). My view: instrument every block in the trace before you upgrade any of them.
  2. Do we need all five? A single-turn classifier needs a model and instructions. A workflow with tools needs no agent state. Adding a block you do not need adds a place to fail. My view: start with model, tools and instructions, add memory when the task outlives one context, add explicit state when runs are long enough to crash.
  3. Who owns each block? In a team, the prompt is often owned by product, the tools by backend engineers, the model choice by a platform team, and memory by nobody. My view: give each block an owner and an eval, or the unowned one will be where the incidents come from.
  4. What is the unit of change? A change to a tool description changes the model's behaviour as surely as a change to the prompt. My view: version all five blocks together, as one release, and test them together.

2.2 The model: choosing the brain

The model is the only block that reasons, and it is also the only block you did not write. You choose it, you cannot debug it, and you can only measure it. Choosing well therefore means knowing what to measure.

What matters for an agent, as opposed to a chat

A model that is excellent at writing essays can be mediocre inside a loop. The agent asks it a narrower question many times over: read this growing context, pick one of these tools, fill in these arguments exactly, or decide that no tool is needed. The factors that matter are different from the ones on a general benchmark.

FactorWhat it means for an agentHow to measure it on your task
Capability on the taskcan it do the hard reasoning steps at all (plan, diagnose, synthesise)?a labelled set of the hard steps, scored by a rubric or a checker
Tool-call reliabilitydoes it pick the right tool, with valid arguments, when it should, and no tool when it should not?tool-selection accuracy, argument accuracy and refusal rate on a labelled set (below)
Latencythe loop multiplies it: ten steps at 3 s is 30 s before any tool timep50 and p95 per call, measured with your real prompt length
Costthe loop multiplies it too, and the context grows every stepcost per task over a sample of real runs, not per call
Context windowcan one run fit, or will you need compaction (Section 2.7)?the largest context your traces reach, with headroom
Structured-output supportcan it be made to emit valid JSON every time, so your parser never breaks?parse-failure rate over a thousand calls
Open-weight versus APIcontrol over data, hosting and version pinning against convenience and the frontier of capabilitya decision about data residency, cost at volume and operations, not a benchmark
The usual shape of the trade-off (a sketch, not a measurement: the table in Section 2.2 says what to measure)large API modelmid-size modelsmall open-weight modelcapability on the taskcapability on the task, large API model: highhighcapability on the task, mid-size model: mediummediumcapability on the task, small open-weight model: lowlowtool-call reliabilitytool-call reliability, large API model: highhightool-call reliability, mid-size model: mediummediumtool-call reliability, small open-weight model: lowlowspeed (low latency)speed (low latency), large API model: lowlowspeed (low latency), mid-size model: mediummediumspeed (low latency), small open-weight model: highhighlow cost per tasklow cost per task, large API model: lowlowlow cost per task, mid-size model: mediummediumlow cost per task, small open-weight model: highhighcontext windowcontext window, large API model: highhighcontext window, mid-size model: mediummediumcontext window, small open-weight model: mediummediumstructured-output supportstructured-output support, large API model: highhighstructured-output support, mid-size model: highhighstructured-output support, small open-weight model: mediummediumcontrol: weights, data, hostingcontrol: weights, data, hosting, large API model: lowlowcontrol: weights, data, hosting, mid-size model: lowlowcontrol: weights, data, hosting, small open-weight model: highhighRead the rows as questions for your eval set, not as a ranking: an agent of ten routine steps and one hard stepmay want the third column for the routine steps and the first for the hard one (the router, below).
The usual shape of the trade-off between three classes of model, sketched qualitatively. Capability and tool-call reliability tend to rise with size; speed, cost and control fall. The sketch is not a measurement: the only numbers that matter are the ones from a per-step evaluation on your own tasks.

What the research says about tool calling

The ability to call a function correctly is a trained skill, and it varies more between models than general fluency does. The Gorilla paper from Berkeley, in May 2023, was the first to measure it carefully.

The same group turned the measurement into a public benchmark, the Berkeley Function-Calling Leaderboard.

The router: a small model for routine steps, a large one for hard ones

Most steps of most agents are routine. Reading a tool result and deciding to call the obvious next tool, extracting a field, reformatting an answer: a small model does these as well as a large one, several times faster and at a fraction of the cost. A few steps are hard: making the plan, recovering from an unexpected error, choosing between two tools whose purposes overlap. The router pattern sends each step to the model that fits it.

A model router inside the loopnext stepfrom the looproutera rule, a classifieror the small model itselfsmall modelroutine: format, extract,pick an obvious toollarge modelhard: plan, recover,ambiguous tool choiceroutinehardtoolescalate on a failedcheck or low confidencepros: most steps are cheap and fast; the large model is paid for only where it changes the outcomecons: two models to evaluate and version; a wrong route is a silent quality loss; the router can be the weak stepto decide: label a per-step eval set routine or hard, run both models on it, route to the small onewherever it matches the large one
A router inside the loop sends routine steps (formatting, extraction, an obvious tool choice) to a small model and hard steps (planning, recovery, ambiguous tool choice) to a large one, with an escalation path when the small model fails a check. A per-step evaluation set, labelled routine or hard, decides where the line goes.
One model for every stepRouter: small for routine, large for hard
Prosone thing to evaluate, version and reason about; no routing errorsmost steps cheap and fast; the large model is paid for only where it changes the outcome
Consevery step pays the price of the hardest steptwo models to evaluate and version; a wrong route is a silent quality loss; the router itself can be the weak step
When to pick itearly, always; and whenever the step mix is mostly hardwhen traces show that most steps are routine and cost or latency is what blocks shipping
How to tell it was rightthe per-step eval passes at the required ratethe per-step eval passes at the same rate as the single large model, at lower cost; the escalation rate is low and stable
Production failurecost or latency too high to servethe small model drifts after a version change and the router does not notice; escalations spike after a prompt edit

Discussion

  1. Should we pick the model first or the tools first? The model is the most visible choice, but tool-call reliability depends on the tool set. My view: design the tools and the eval set first, then let the eval choose the model; otherwise you choose a model for tools you will redesign.
  2. Is a bigger model ever the wrong fix? Yes: when the failure is a vague tool description or a missing stop condition, a bigger model reads the same ambiguity more cleverly and still guesses. My view: upgrade the model only after the trace shows the step was reasoned well and still wrong.
  3. How do we handle model version changes? Vendors retire and replace models; open weights do not change under you but need operations. My view: pin versions explicitly, treat an upgrade as a release with the full eval, and never let "latest" into production.
  4. Open weights or API? The honest answer depends on volume, data rules and the team's operations capacity. My view: start on an API to learn what the task needs; move the routine steps to an open-weight model only when the per-step eval proves parity and the volume pays for the operations.

2.3 Tools: giving the model hands

A model can only produce text. Everything an agent does happens through a tool, and the mechanism by which text becomes an action is what vendors call function calling (or tool use). Because the model never runs anything itself, tools are the boundary where your code, your permissions and your safety checks live. They are also, in the experience of every team that has published about it, the block most worth engineering carefully.

What a tool is and how the round trip works

your code (the loop)the model (API)1request: system prompt + tool definitions + user turntools = [{name, description, input_schema}]user: "Find 45 min for Priya and Tom..."2reply: a structured call, stop_reason = tool_use{type: tool_use, id: call_01, name: calendar_find_free_slots, input: {attendee_emails: [...], ...}}3your code validates the arguments and runs the functionresult = find_free_slots(**input)the model never executes anything4request: the result goes back as a tool_result turn{type: tool_result, tool_use_id: call_01, content: "{slots: [...], time_zone: ...}"}5reply: text (or another tool call), stop_reason = end_turn"Three slots fit (none before 10:00) ..."steps 1 to 4 repeat while the model keeps asking for tools; the whole exchange is the context the next call sees
One function-calling round trip. The request carries the system prompt, the tool definitions and the user turn; the model replies with a structured tool call and a stop reason that says so; your code validates the arguments and runs the function; the result goes back as a tool-result turn; the model replies with text or with another call. Steps one to four repeat while the model keeps asking for tools, and the whole exchange is the context the next call sees.

The script ch2_tools.py prints the exact messages of one round trip for the meeting-scheduling example that runs through this chapter. The model is a stub (a function that returns what a real model would be expected to return at each point), so the script runs with no API key; the tool definition and the message shapes are real.

Terminal output of ch2_tools.py part one: the tool definitions sent in the request, the full JSON schema of calendar_find_free_slots with attendee_emails, duration_minutes, window_start, window_end and earliest_hour, the user turn, the assistant reply containing a tool_use block with id call_01 and filled-in arguments, the tool_result turn sent back with three slots and a time zone, and the final assistant text listing the three slots

Four things in that output are the whole mechanism. The tool definitions travel in the request, so they cost input tokens on every turn. The model's reply is not text but a tool_use block with an id, a name and an input object that matches the schema. Your code, not the model, runs find_free_slots, and it is your code that could refuse, log, rate-limit or ask a human first. The result goes back as a tool_result tied to the id, and the model's next reply is ordinary text because it decided it was done.

Where tool use came from

Chapter 1 introduced Toolformer as the paper that showed tool use can be learned rather than prompted. Its first figure is worth a second look here, because it shows what a tool call is from the model's point of view: a piece of text, in a fixed format, inserted exactly where the information is needed.

Scale came next. The ToolLLM paper (July 2023) asked what happens when the tool set is not four tools but thousands.

How to design tools a model can use

The most useful engineering guidance on this is Anthropic's September 2025 post "Writing effective tools for agents", which came out of optimising the company's own internal tools against evaluations. Its five principles are a checklist; this section reads them one at a time and adds the trade-offs.

1. Choose which tools to build, and do not build the rest.

2. Namespace the tools.

3. Return meaningful context, not raw payloads.

4. Make responses token-efficient.

5. Prompt-engineer the descriptions.

Anthropic's December 2024 post "Building effective agents" made the same point from the other direction a year earlier, in an appendix on what it calls the agent-computer interface.

a tool the model will misusea tool the model can usenameget_datadescriptionGets data.argumentsq: string (what goes in q?)returnsthe whole table, 40,000 tokenson error"Error 500"on emptyraises an exceptionside effectsunknown; sometimes writesidempotentno: a retry creates duplicatesnamecalendar_find_free_slotsdescriptionFind slots when ALL attendees are free.Use before booking; [] if none are free.argumentsattendee_emails[], duration_minutes,window_start, window_end, earliest_hourreturnsat most 10 slots + the time zoneon error"window_end is before window_start:swap them or widen the window"side effectsnone (read-only, annotated)idempotentyes: safe to retrythe model has to guess what q is, reads 40k tokensto find one row, and cannot tell a failure froman empty answer; a retry books twicename says the service and the verb; arguments arenamed so they cannot be confused; the result fits inthe context; the error says what to do next
A tool the model will misuse next to a tool it can use. The good tool has a namespaced name, a description that says when to use it and what it returns, unambiguous argument names, a bounded return size, an error message that says what to do next, no side effects and safe retries. The bad tool forces the model to guess the argument, read forty thousand tokens to find one row, and cannot tell a failure from an empty answer.
PropertyA tool the model will misuseA tool the model can useWhy it matters in the loop
Nameget_data, fetch, do_calendarcalendar_find_free_slots (service, resource, verb)the name is read first when choosing
Description"Gets data."what it does, when to use it, what it returns, what it returns when there is nothingthe only training the model gets on the tool
Argument countone vague q: string, or twelve optional flagsthree to six named, typed arguments with defaultseach argument is a place to be wrong
Argument namesuser, date, iduser_id, window_start (ISO date), event_idambiguity becomes a wrong value
Return sizethe whole tableat most N items, paginated or filtered, with a capevery token returned is re-read every later turn
Error messages"Error 500", a stack trace"window_end is before window_start: swap them or widen the window"the model can only recover from errors it can read
Empty resultraisesreturns an empty list and says soa model that sees an exception may retry forever
Idempotencya retry creates a second eventsafe to retry, or takes an idempotency keythe loop will retry
Side effectsundocumenteddeclared (read-only, destructive, external) and checked by the looppermissions live at this boundary

MCP: a standard way to expose tools

Until late 2024 every team wired its tools into its agent by hand, and every tool vendor shipped a different integration for every agent product. The Model Context Protocol (MCP), released by Anthropic in November 2024 and since adopted widely, standardises the wire: a server exposes tools (and two other things) over a protocol; any host that speaks MCP can use them.

The Model Context Protocol: one host, one client per server, three primitives per serverMCP host (the AI application: an IDE, a desktop app, your agent)modelthe agent loopMCP client AMCP client Bthe host lists tools from every client andputs their definitions into the model's contextMCP server: filesystem (local, stdio)tools: read_file, search_filesresources: file contentsprompts: "summarise this folder"MCP server: issue tracker (remote, HTTP)tools: issues_search, issue_createresources: an issue, a projectprompts: "triage this issue"both links speak JSON-RPC: tools/list to discover, tools/call to runhost: coordinates the clients and owns the model · client: one connection to one serverserver: provides tools (actions), resources (data) and prompts (templates)
The participants of the Model Context Protocol as its documentation describes them. The host (your agent, an IDE, a desktop application) creates one client per server; each server exposes tools, resources and prompts over JSON-RPC, locally over standard input and output or remotely over HTTP; the host lists the tools from every client and puts their definitions into the model's context.
Hand-written tool wrappersMCP servers
Prosfull control of names, descriptions, return shapes and error text; no protocol overhead; trivial to testreuse of hundreds of existing servers; one integration per host, not per tool; dynamic discovery; standard annotations for destructive or open-world tools
Consevery tool is bespoke work; every agent product needs its own integrationyou inherit the server author's names, descriptions and return sizes; many servers wrap endpoints one-to-one, exactly the anti-pattern above; tool definitions from many servers can flood the context
When to pick ita small, stable tool set you own, where every description is tuned on your evalmany systems to connect, especially third-party ones; or you are shipping tools for other people's agents
How to tell it was righttool-selection and argument accuracy on your eval; return sizes within budgetthe same metrics, measured with the servers connected; plus the context cost of the loaded definitions
Production failurea wrapper silently drifts from the API it wrapsa server update renames a tool or changes a return shape and your prompt's tool guidance is now wrong; a server with sixty tools pushes your definitions past the attention budget

The context cost of many servers is real and measured. Anthropic's November 2025 post on code execution with MCP describes an agent that, instead of loading every tool definition, writes code that discovers tools as files and calls them from a sandbox.

Failure modes

Every one of these appears in traces of real systems, and each has a cheap detector.

FailureWhat it looks like in the traceWhy it happensDetector
Too many toolsthe model picks a plausible but wrong tool; long deliberation; more steps than neededdefinitions compete for attention; many are never relevanttool-selection accuracy falls as tools are added; definition tokens per turn
Overlapping toolssearch and find and lookup chosen interchangeably, with different resultsno clear boundary in names or descriptionsconfusion matrix between tools on the eval set
Huge return payloadsone observation of tens of thousands of tokens; later steps lose earlier factsthe tool returns what the API returnsobservation size per step; p95 of tool result tokens
Success on failurethe tool returns 200 OK with an error string inside; the model carries on as if it workedthe wrapper does not inspect the bodytool results containing error words with a success status
Ambiguous argumentsdate: "next Monday", user: "Priya" where an id was neededthe schema allows free text where it should notargument validation failures; downstream "not found" errors
Non-idempotent retriestwo events, two emails, two ticketsthe loop retried a write after a timeoutduplicate detection on side effects
Missing "no tool" optionthe model calls a tool for a question it could answer, or calls a wrong one when none fitsnothing in the instructions says when not to callrefusal-rate cases in the eval

How to evaluate a tool set

Three numbers, all from a labelled set of real requests.

  1. Tool-selection accuracy: did the model pick the labelled tool (or correctly pick none)? Report it per tool as well as overall; one badly described tool can hide behind a good average.
  2. Argument correctness: schema-valid and semantically right. Schema validity is free; semantic correctness needs the label to contain the expected arguments.
  3. Trajectory length: how many tool calls did the task take against the labelled minimum? Extra calls mean the tools do not match how the task decomposes, or the returns do not contain what the next step needs.

The second half of ch2_tools.py shows the mechanism of the first metric with a deliberately simple stub: a chooser that picks the tool whose name and description share the most words with the request.

Terminal output of ch2_tools.py part two: ten labelled requests run against a set of five bare endpoint wrappers named get_calendar, post_calendar_event, get_crm_record, post_mail and get_search, with seven of ten wrong and 3 of 10 correct, then the same ten requests against five described tools named calendar_find_free_slots, calendar_create_event, crm_get_customer_context, email_send and docs_search, all ten correct; the last line reads selection accuracy wrappers 30 percent, clear 100 percent, a stub, not a model, it shows the mechanism, not a benchmark

Read the 30% against 100% for what it is: the output of a word-overlap stub, not a model's score. What it demonstrates is why descriptions matter. The wrapper set (GET /calendar, POST /mail) shares almost no vocabulary with how people ask for things, so the chooser has nothing to match and falls back on the first tool. The described set names the service, the verb and the words a user would use. A real model is far better than a word count at bridging that gap, but it bridges it with the same material, and the gap is what your eval measures.

Discussion

  1. One tool per endpoint or one tool per task? Endpoint wrappers are quick to generate and map to documentation you already have; task-shaped tools need design and an eval. My view: generate the wrappers to learn the domain, then replace the ones on the hot path with consolidated, task-shaped tools, and measure the trajectory length drop.
  2. How many tools is too many? There is no fixed number; the symptom is selection accuracy falling and deliberation rising. My view: if a human engineer cannot say which tool to use in a given situation without thinking, the model cannot either; past a dozen or two, load tools per step rather than all at once.
  3. MCP or in-house? MCP buys reach and costs control. My view: use MCP for everything you do not own and would otherwise integrate badly; write your own tools for the five that decide whether your agent works, and tune their descriptions on your eval.
  4. Should tools be forgiving or strict? A forgiving tool (accepts "next Monday") raises step success now; a strict one (requires an ISO date) raises it permanently by making the ambiguity visible. My view: strict schemas with helpful error messages; the model learns from the error on the next turn, and the trace shows you where the schema should change.
  5. Who approves a destructive tool? The annotation says destructive; something must act on it. My view: the loop, not the model, checks annotations and routes destructive calls to the approval step of Chapter 1's level 2; never trust the model to ask permission on its own.

2.4 Instructions: the system prompt as architecture

The system prompt is usually written last, by whoever is closest to the demo, and then grows by accretion as each incident adds a sentence. Treated that way it becomes the least reliable block in the system. Treated as architecture, it is the cheapest place to raise per-step reliability, because every word in it is read on every turn.

The anatomy

The anatomy of an agent's instructions, and the smell when a section is missingrole and goal"You schedule meetings for the sales team. Your job ends when the invite is sent."missing: the agent does adjacent tasks nobody asked forhard constraints"Never book before 10:00. Never email outside the company. Ask before deleting."missing: the guardrail lives only in people's headstool guidance"Use find_free_slots before create_event. Prefer one search with filters."missing: the right tool, called in the wrong orderoutput format"Reply with a one-line summary and the event id. No markdown."missing: a parser that breaks on every third runstop conditions"Stop when the invite is sent, after 8 steps, or if a tool fails twice."missing: loops that never end, budgets that run outcanonical examplestwo or three short, diverse transcripts of a good runmissing: edge cases listed as rules instead of shownorder matters less than clarity; keep every section short enough that a new colleague would read it
The six sections of an agent's system prompt, each with an example line and the production smell that appears when it is missing: role and goal (the agent does adjacent tasks nobody asked for), hard constraints (the guardrail lives only in people's heads), tool guidance (the right tool called in the wrong order), output format (a parser that breaks every third run), stop conditions (loops that never end) and canonical examples (edge cases listed as rules instead of shown).

Role and goal. One or two sentences: what the agent is, who it serves, and what finishing looks like. The goal matters more than the role; "your job ends when the invitation has been sent" prevents more drift than three paragraphs of persona.

Hard constraints. The rules that must hold whatever the user says: never book before 10:00, never email outside the company, ask before deleting. These are the first line of guardrails (Chapter 9), and they belong in the prompt and in code, because the prompt is a request and code is a guarantee.

Tool guidance. When to use which tool and in what order; when to use none. The tool descriptions say what each tool does; the prompt says how they fit together for this task ("find free slots before creating an event; prefer one filtered search to many broad ones").

Output format. What the final answer must look like, so that the code that reads it never guesses. For an agent, the final answer is often a structured object, and the format section is a schema.

Stop conditions. When the agent is done, how many steps it may take, and what to do when a tool fails repeatedly. Chapter 1 showed "unaware of termination conditions" as one of the most common multi-agent failures; the fix starts here.

Canonical examples. Two or three short transcripts of a good run, chosen for diversity rather than coverage.

What the vendors advise

OpenAI's practical guide to building agents (April 2025) devotes a page to instructions and lists four practices.

Anthropic's September 2025 post on context engineering describes the balance to strike.

Long or short instructions

Long, detailed instructionsShort, minimal instructions
Prosmore control; edge cases handled; less dependence on the model's defaultsfewer tokens on every turn; fewer internal contradictions; easier to test and reason about; lets a capable model use its judgement
Constokens paid on every step; rules conflict with each other and with tool descriptions; later rules get lost in the middle (Section 2.7); nobody knows which sentence does whatrelies on the model doing the sensible thing; edge cases surface in production
When to pick ita regulated task with hard rules; a weaker model that needs them; a stable task whose edge cases are knownearly, always; a capable model; a task where the right action depends on judgement
How to tell it was rightthe eval passes and removing any paragraph makes it fail (otherwise the paragraph is dead weight)the eval passes; the failure analysis finds no cluster that a sentence would fix
Production failurea new rule added after an incident silently contradicts an old one; cost creepa long tail of unhandled cases, each handled by a hot-fix sentence until the prompt is long anyway

Prompt versioning and testing as code

The practical rules are the ones you already use for code. Store the prompt as a file, not a string buried in a function. Review changes as diffs. Run the eval set in CI on every change, and gate the release on it (Chapter 7 builds that gate). Log the prompt version in every trace, so a regression can be traced to the edit that caused it. And remove sentences as readily as you add them: the eval that justified a sentence should be re-run without it from time to time, because models change and the sentence may now be noise.

Instruction smells

SmellExampleWhat it causesFix
Rules that contradict"always confirm before booking" and "complete the task without asking"random behaviour per runone rule, with the exception stated
Hidden stopno sentence says when the task is donethe agent keeps improving its answer until the budget runs outan explicit done condition and a step budget
Tool guidance in the wrong placethe prompt describes what find_free_slots returnsthe description and the prompt drift apartdescribe the tool in the tool; sequence the tools in the prompt
Edge cases as rulestwenty "if the user says X then Y" linesbrittle; the twenty-first case breaks itthree canonical examples that show the judgement
Shouting"NEVER EVER", "CRITICAL", "YOU MUST"emphasis inflation; everything is critical so nothing isplain sentences; the hard constraints also enforced in code
Negatives without alternatives"do not use markdown"the model does not know what to do instead"reply in plain sentences"
Persona over purposethree paragraphs of character, one line of goaldrift and verbosityone line of role, a clear goal, a clear finish
Untested editsa sentence added after an incident, no eval runfixes one case, breaks threeprompt in version control, eval in CI

Discussion

  1. How much should the prompt know about the tools? Too little and the model uses them in the wrong order; too much and the prompt duplicates the descriptions and drifts. My view: tool descriptions say what; the prompt says when and in what order, in one short section.
  2. Should hard constraints be in the prompt if they are also in code? Both. The prompt tells the model what to avoid so it does not waste steps proposing it; the code stops it when the prompt fails. My view: never only in the prompt.
  3. Who edits the prompt? Product people understand the task; engineers understand the loop. My view: anyone can propose, in a pull request, with the eval run attached; nobody edits in production.
  4. Can we trust a vendor's prompting guide? Their advice is tuned for their models and changes with each release. My view: adopt the structure (sections, altitude, canonical examples), and test the specifics on your eval.

2.5 Memory: what the agent remembers and for how long

A model remembers nothing between calls. Everything an agent "remembers" is engineering: something your code kept and put back in front of the model. The design questions are what to keep, where, and for how long, and the right answers differ by kind of information.

Three places memory lives

where memory liveswhat kind of thing is rememberedshort-term: the context windowthis turn; the model sees it directly;costs tokens on every call; gone at the endworking: scratchpad and planthis run; written by the agent as notes;survives compaction; small by designlong-term: a store outsideacross runs; written and retrieved by key,search or embedding; cheap to hold, costly to get wrongwriteretrievesummarisepull inepisodicwhat happened"on Monday the invite bounced, Tom was out";past runs, traces, reflections (Reflexion)semanticfacts about the world and the user"Priya is in Dublin"; "the user never meetsbefore 10:00"; a profile, a knowledge baseproceduralhow to do things"always check the room before booking"; theinstructions, learned rules, tool recipes
Left: the three places memory lives. The context window is seen directly by the model and paid for on every turn; the working memory is a scratchpad the agent writes for itself during a run; the long-term store lives outside the context, across runs, and is written to and retrieved from. Right: the three kinds of memory, episodic (what happened), semantic (facts) and procedural (how to), each with an example.

Three kinds of thing remembered

The second axis comes from cognitive science and has been adopted by agent frameworks because it tells you what to store where.

Memory typeCostStalenessPrivacyRetrieval errorsTypical store
Short-term (context)highest: paid on every turn, grows every stepnone: it is the presentcontained in the runnone: the model sees all of it, though attention fades (Section 2.7)the message list
Working (scratchpad)low: a few hundred tokens per turnlow: rewritten by the agent as it goescontained in the runlow: small and structureda notes field, a plan object, a notes file
Long-term episodicstorage cheap; retrieval has a token cost per itemhigh: events age; outcomes get supersededhighest: records of real people's interactions, with retention obligationsretrieving an irrelevant or outdated episode misleads the agentvector store over summaries, logs with embeddings
Long-term semanticstorage cheap; retrieval modestmedium: facts change and must be corrected, not appendedmedium: user profiles are personal dataretrieving a contradicted fact; two versions of the same factkey-value by entity, a profile table, a knowledge base
Long-term procedurallow: a few ruleslow to medium: rules outlive the reason for themlowa learned rule that was wrong propagates to every runthe prompt; a rules file the agent may edit under review

The research: a memory stream, and an operating system

The Generative Agents paper (April 2023) is where the modern memory architecture for agents was written down.

The paper's retrieval function is a small formula worth knowing, because every "which memories should go into the prompt" decision ends up looking like it.

Generative Agents: score = recency + importance + relevance, then keep the top ones that fitquery: "what should Klaus do this afternoon?" (illustrative values, chosen to show the mechanism)recencyimportancerelevancescoreIsabella is planning a party on the 14thIsabella is planning a party on the 14th: recency 0.550.55Isabella is planning a party on the 14th: importance 0.850.85Isabella is planning a party on the 14th: relevance 0.950.952.35met Klaus at Hobbs Cafe, talked about researchmet Klaus at Hobbs Cafe, talked about research: recency 0.700.70met Klaus at Hobbs Cafe, talked about research: importance 0.500.50met Klaus at Hobbs Cafe, talked about research: relevance 0.600.601.80the refrigerator is emptythe refrigerator is empty: recency 0.950.95the refrigerator is empty: importance 0.200.20the refrigerator is empty: relevance 0.100.101.25ate breakfast at 8 in the roomate breakfast at 8 in the room: recency 0.900.90ate breakfast at 8 in the room: importance 0.100.10ate breakfast at 8 in the room: relevance 0.150.151.15recency decays with the time since the memory was last used (0.995 per game hour in the paper);importance is a score the model gave the memory when it was stored; relevance is the embeddingsimilarity to the query. All three weights are 1 in the paper. The two highlighted memories go into the prompt.
The retrieval function of Generative Agents on four illustrative memories. Each gets a recency score, an importance score assigned when it was stored, and a relevance score to the current situation; the paper sums them with equal weights, and the top-ranked memories that fit in the context go into the prompt. The two highlighted memories are selected; the recent but trivial ones are not. Values are illustrative.

Six months later, MemGPT (October 2023) gave the problem its most useful analogy.

MemGPT: the context window as main memory, external stores as disk, the model as its own operating systemmain context (fixed size: the prompt tokens)system instructions (how to use the memory functions)working context: facts the model chose to keep"birthday is February 7", "boyfriend named James"FIFO queue: the most recent messagesoldest ones are evicted (and summarised) firstalert when the queue nears the limit: "memory pressure"external context (unbounded: a database)recall storagethe full message history, searchablerecall_storage.search("six flags")archival storagedocuments and notes the model filed awayarchival_memory.insert(...)only what is paged in is ever seen by the modelpage outpage inthe paging is done by function calls the model itself makes when it sees the alert; the loop only runs them
The operating-system analogy of MemGPT. The fixed-size context window is main memory, holding the system instructions, a working context of facts the model chose to keep, and a queue of recent messages; recall storage (the full searchable history) and archival storage (filed documents and notes) are disk. The model pages information in and out by calling memory functions when the loop warns it of memory pressure.

Memory in production: compaction, notes, sub-agents

Anthropic's context-engineering post names three techniques its own long-running agents use, and they map exactly onto the ladder.

A toy memory in fifty lines

The script ch2_memory.py implements the ladder in miniature: a history (short-term), a scratchpad (working), a key-value store (long-term), and a summariser stub that compacts the history every four steps by keeping the first ninety characters of each old message (a real system would ask the model to summarise). It replays twelve steps of the meeting task and prints the approximate context size per step, with and without compaction. Token counts use the rule of thumb of about four characters per token; they are an approximation, not a tokenizer.

python
# code/agents/ch2_memory.py (abridged; the summariser is a stub and the token count is chars // 4)
class Memory:
    def __init__(self, compact_every=4, keep_last=2):
        self.history, self.scratch, self.store = [], [], {}  # short-term, working, long-term

    def remember(self, key, value): self.store[key] = value             # long-term: survives compaction
    def recall(self, query): return {k: v for k, v in self.store.items() if any(w in k for w in query.split())}
    def note(self, line): self.scratch.append(line)                      # working memory: the running plan

    def add(self, role, content, step):
        self.history.append({'role': role, 'content': content})
        if step % self.compact_every == 0:                               # compaction: old turns -> one summary
            old, recent = self.history[:-self.keep_last], self.history[-self.keep_last:]
            summary = ' '.join(m['content'][:90] for m in old if m['role'] != 'summary')
            self.history = [{'role': 'summary', 'content': summary}] + recent

    def context(self, query):                                            # what the model would see this turn
        return '\n'.join(['SYSTEM: ...'] + [f'MEMORY {k}: {v}' for k, v in self.recall(query).items()]
                         + ['PLAN: ' + ' | '.join(self.scratch)] + [f'{m["role"]}: {m["content"]}' for m in self.history])

Terminal output of ch2_memory.py: without compaction the context grows from 79 to 489 approximate tokens over twelve steps, 3,660 in total; with compaction every four steps it peaks at 349 and ends at 296, 2,772 in total; then the compacted context at step 12 showing the system line, two MEMORY lines including the preference never before 10:00, a PLAN line of notes, a SUMMARY line, and the last two turns; finally recall of the word meetings returns the stored preference, with the note that the user turn stating the rule was summarised away at step 4 but survives in the long-term store

0100200300400500123456789101112without compaction, step 1: 79without compaction, step 2: 165without compaction, step 3: 192without compaction, step 4: 224without compaction, step 5: 246without compaction, step 6: 296without compaction, step 7: 321without compaction, step 8: 361without compaction, step 9: 390without compaction, step 10: 434without compaction, step 11: 463without compaction, step 12: 489with compaction, step 1: 79with compaction, step 2: 165with compaction, step 3: 192with compaction, step 4: 164with compaction, step 5: 187with compaction, step 6: 237with compaction, step 7: 261with compaction, step 8: 246with compaction, step 9: 276with compaction, step 10: 320with compaction, step 11: 349with compaction, step 12: 296without compaction: 489with compaction: 296step of the toy runapprox. tokens in contexttoy run, 4 chars per token: 3,660 tokens total without 2,772 with (every 4 steps)
Approximate context size per step of the twelve-step toy run, with and without compaction every four steps. The compacted run drops at steps 4, 8 and 12 and ends about forty percent smaller; over twelve steps it sends about a quarter fewer tokens. The numbers are from the four-characters-per-token approximation in ch2_memory.py.

Two things to notice. The saving is modest at twelve steps (3,660 against 2,772 tokens) because the run is short; Section 2.7 shows what happens at thirty. More important is the last two lines of the output: the user's rule "never before 10:00" was stated at step 1, compacted away at step 4, and is still available at step 12 because it was also written to the long-term store. Without that write, the agent would have booked Tue 09:00 at step 10 with a clear conscience. The summariser kept the first ninety characters of each message, which happened to include the rule in this run; a real summariser might not. Memory that matters must not depend on what a summary happens to keep.

How to evaluate memory

MetricWhat it measuresHow
Recall after N turnscan the agent still use a fact stated N turns ago?plant facts at step k, ask at step k+N, score the answer; sweep N past the compaction boundary
Contradiction ratedoes the agent act against a stored fact or against something it said earlier?a checker over the trace (booked before 10:00 when the store says never)
Retrieval precisionof the memories put into the context, how many were used or relevant?label retrieved items; or ablate them and see whether the answer changes
Cost of context per stepwhat does memory add to each call?tokens of retrieved memory and notes per step, from the trace
Stalenesshow often does a retrieved fact turn out to be superseded?compare retrieved facts to the current truth where you have it

Discussion

  1. What is worth storing long-term? Everything is cheap to store and expensive to retrieve wrongly. My view: store semantic facts about the user by key, episodic summaries only where a past outcome changes a future decision, and procedural rules only under human review.
  2. Should the model decide what to remember, or should code? MemGPT lets the model decide; a profile table lets code decide. My view: let the model propose, and have code (or a reviewer) accept, for anything that will influence future runs for other people.
  3. How aggressively should we compact? Aggressive compaction saves money and loses facts. My view: tune for recall first (plant facts and check they survive), then trim; clear tool results before you summarise anything a user said.
  4. What do we owe the user about their memory? Episodic and semantic memory are records of a person. My view: show them what is stored, let them delete it, and set retention before launch, because the first deletion request will come after.
  5. Is a vector store the default? It is the default in tutorials, not in production. My view: a key-value profile and a notes file cover most agents; add similarity search when you have many unstructured memories and a measured retrieval problem.

2.6 State: where the loop is

Memory is what the agent knows. State is where it is: which step, what it planned to do, which call is in flight, how much budget is left, what has gone wrong. Confusing the two produces agents that cannot be resumed, retried or inspected. Keeping state explicitly produces agents that can.

The state object of one run, checkpointed after every step1after step 1step: 1 of 8plan: slots → book → invitepending: find_free_slotsbudget: 7 steps, 38k tokerrors: 0status: runningcheckpoint2after step 2step: 2 of 8plan: slots ✓ → bookpending: nonelast: 3 slots foundbudget: 6 steps, 31k tokstatus: runningcheckpoint3after step 3step: 3 of 8plan: book ✓ → invitepending: send_emailevent: EVT-1187 (Mon)budget: 5 steps, 24k tokstatus: runningcheckpoint4after step 4step: 4 of 8plan: invite ✗ → movepending: noneerrors: 1 (Tom out Mon)budget: 4 steps, 17k tokstatus: needs_retrycheckpointmemory says what the agent knows; state says where it is. A crash after step 3 resumes from checkpoint 3with the pending call still recorded, so the retry is deliberate; the email tool must be idempotentor the second attempt invites everyone twice.
The state object of a meeting-scheduling run over four steps: the current step, the plan with progress marks, the pending tool call, the budget left, the errors so far and the status. Each step ends with a checkpoint, so a crashed run resumes from the last one instead of starting again; the pending call is recorded so a retry is deliberate, and the email tool must be idempotent or the second attempt invites everyone twice.

Why does the pending call matter? Suppose the process dies between sending the email and recording the result. On resume, the state says "pending: send_email". The loop now has a choice: run it again (safe only if the tool is idempotent), check whether it happened (if the tool offers a way), or ask a person. Without the pending field, the loop does not even know there is a question. Chapter 1's account of Anthropic's research system mentioned exactly this: the team added "the ability to resume from where an error occurred rather than restart", and used deployments that let in-flight agents finish on the old version. Both are state engineering.

The state machine view

An agent loop can be drawn as a graph: nodes are steps (call the model, run a tool, ask for approval, summarise), edges are transitions, and the state object travels along the edges. Frameworks such as LangGraph make this explicit, and attach persistence to it.

Free-running loopExplicit state machine (graph)
Prossimplest possible code; the model has full freedom; new tools need no graph changeevery possible transition is visible and testable; approval steps, retries and compaction are nodes, not special cases; checkpointing falls out of the design
Conshard to resume; hard to insert an approval step; the only budget is a step counter; behaviour is wherever the model took itmore code and more concepts; the graph can over-constrain a capable model; frameworks bring their own abstractions and upgrades
When to pick itprototypes; short runs (a handful of steps) where restarting is cheapruns long enough to crash or to need approval; regulated actions; anything that must resume
How to tell it was rightruns finish within budget and nobody needs to resume oneresumed runs complete at the same rate as fresh ones; approval steps are hit exactly when the graph says; no duplicate side effects after retries
Production failurea deploy kills in-flight runs; a timeout retries a write; nobody can say what step a stuck run is onthe graph does not allow the step the task needed; checkpoints grow unbounded; two versions of the graph disagree about a checkpoint's shape

Discussion

  1. Where does the plan live, memory or state? It is both: the plan's content is working memory, its progress is state. My view: keep the plan text in the scratchpad and the progress marks in the state object, so compaction never erases where you are.
  2. Framework or your own loop? A graph framework gives you checkpointing and approval nodes on day one and a dependency on day two. My view: write the free loop first to understand the task, then adopt a graph (your own or a framework's) the moment you need to resume or to pause for a human.
  3. How much should the model see of the state? Showing it the budget ("3 steps left") improves stopping; showing it the whole state invites it to reason about plumbing. My view: show the budget and the errors, hide the rest.
  4. Retry or ask? After a crash with a pending write, retrying is fast and asking is safe. My view: retry only idempotent calls automatically; route everything else to a person, and make the trace show which happened.

2.7 Context engineering: the art of what to put in the window

Every block in this chapter ends up as tokens in one place: the context the model reads on this turn. The model sees nothing else. It does not see your database, your notes file, your tool implementations or your intentions; it sees the window. Context engineering is deciding what goes into that window at each step, and it is the discipline that ties the five blocks together.

The budget

What one turn sends to the model at step 10 of a run: 17,100 tokens (planning numbers from ch2_math.py)system instructions: 1,500 tokenstool definitions (20 x 150): 3,000 tokensretrieved memory: 1,000 tokenshistory so far (step 10): 10,800 tokenshistory 10,800current observation: 800 tokensinstructions 1,500tool definitions 3,000memory 1,000observation 800the history is already 63% of the turn and grows by about 1,200 tokens per step; by step 30 the samerun sends 40,300 tokens per turn without compaction (Section 2.7). The window's hard limit is far abovethis; the practical limit is attention (Section 2.7).
The context of one turn as a budget bar at step 10 of a run, using the planning numbers in ch2_math.py: 1,500 tokens of instructions, 3,000 of tool definitions (twenty tools at about 150 each), 1,000 of retrieved memory, 10,800 of history and 800 for the current observation, 17,100 in all. The history already dominates and grows every step; the fixed parts are paid for on every call.

The numbers in the figure are planning numbers, not measurements: the point is the shape. At step 10 the history is already 63% of the turn, and it grows by about 1,200 tokens per step (a model turn of about 400 tokens plus an observation of about 800). The instructions and tool definitions, 4,500 tokens here, are paid on every single call whether or not the step needs them. The retrieved memory is small only because somebody made it so.

Why the window's hard limit is not the real limit

Models with very long windows exist, and the temptation is to treat the limit as the budget. Two findings say otherwise. The first is from Stanford and Berkeley, in July 2023.

The second finding is the one Anthropic's post calls context rot: as the number of tokens in the window increases, the model's ability to recall information from it decreases, so that context "must be treated as a finite resource with diminishing marginal returns". The post's explanation is architectural: attention relates every token to every other, and models have seen far more short sequences than long ones in training, so precision at long range is a "performance gradient rather than a hard cliff".

The arithmetic of a thirty-step run

The script ch2_math.py runs the budget forward for thirty steps, with and without compaction every ten steps (the history replaced by a 600-token summary plus the last two turns), at the same example price as Chapter 1: $3 per million input tokens and $15 per million output tokens, round numbers for the arithmetic and not a quote from any provider.

Terminal output of ch2_math.py: part one, the budget of one turn at step 10 with instructions 1,500, tool definitions 3,000, retrieved memory 1,000, history 10,800 and observation 800, total 17,100; part two, a table of input tokens per step with and without compaction, 5,500 at step 1 for both, 16,300 at step 10, then 17,500 against 8,500 at step 11 and 40,300 against 19,300 at step 30; totals of 687,000 against 387,000 input tokens, 12,000 output tokens, cost 2.241 against 1.341 dollars, 44 percent fewer input tokens and 40 percent cheaper; part three, a 25,000-token tool result at step 5 re-sent on every later step adds 650,000 input tokens and 1.95 dollars per run, against 20,800 tokens if trimmed to 800

010,00020,00030,00040,000151015202530no compaction, step 1: 5,500no compaction, step 2: 6,700no compaction, step 3: 7,900no compaction, step 4: 9,100no compaction, step 5: 10,300no compaction, step 6: 11,500no compaction, step 7: 12,700no compaction, step 8: 13,900no compaction, step 9: 15,100no compaction, step 10: 16,300no compaction, step 11: 17,500no compaction, step 12: 18,700no compaction, step 13: 19,900no compaction, step 14: 21,100no compaction, step 15: 22,300no compaction, step 16: 23,500no compaction, step 17: 24,700no compaction, step 18: 25,900no compaction, step 19: 27,100no compaction, step 20: 28,300no compaction, step 21: 29,500no compaction, step 22: 30,700no compaction, step 23: 31,900no compaction, step 24: 33,100no compaction, step 25: 34,300no compaction, step 26: 35,500no compaction, step 27: 36,700no compaction, step 28: 37,900no compaction, step 29: 39,100no compaction, step 30: 40,300compaction every 10 steps, step 1: 5,500compaction every 10 steps, step 2: 6,700compaction every 10 steps, step 3: 7,900compaction every 10 steps, step 4: 9,100compaction every 10 steps, step 5: 10,300compaction every 10 steps, step 6: 11,500compaction every 10 steps, step 7: 12,700compaction every 10 steps, step 8: 13,900compaction every 10 steps, step 9: 15,100compaction every 10 steps, step 10: 16,300compaction every 10 steps, step 11: 8,500compaction every 10 steps, step 12: 9,700compaction every 10 steps, step 13: 10,900compaction every 10 steps, step 14: 12,100compaction every 10 steps, step 15: 13,300compaction every 10 steps, step 16: 14,500compaction every 10 steps, step 17: 15,700compaction every 10 steps, step 18: 16,900compaction every 10 steps, step 19: 18,100compaction every 10 steps, step 20: 19,300compaction every 10 steps, step 21: 8,500compaction every 10 steps, step 22: 9,700compaction every 10 steps, step 23: 10,900compaction every 10 steps, step 24: 12,100compaction every 10 steps, step 25: 13,300compaction every 10 steps, step 26: 14,500compaction every 10 steps, step 27: 15,700compaction every 10 steps, step 28: 16,900compaction every 10 steps, step 29: 18,100compaction every 10 steps, step 30: 19,300no compaction: 40,300compaction every 10 steps: 19,300step of the runinput tokens on this stepinput tokens over the run: 687,000 without 387,000 withcost at the example price: $2.24 vs $1.34
Input tokens sent on each step of a thirty-step run. Without compaction the context grows linearly to 40,300 tokens at step 30 and the run sends 687,000 input tokens in all; with compaction every ten steps the context never exceeds 19,300 and the run sends 387,000. At the example price the run costs $2.24 against $1.34, about forty percent less.
Without compactionCompaction every 10 steps
Context at step 15,5005,500
Context at step 1117,5008,500
Context at step 3040,30019,300
Input tokens over the run687,000387,000
Output tokens over the run12,00012,000
Cost at the example price$2.24$1.34

Three readings of the table. First, compaction cut input tokens by 44% and cost by 40% in a thirty-step run, and the saving grows with length because the uncompacted cost is quadratic in the number of steps (each step re-sends everything before it). Second, the largest context halved, from 40,300 to 19,300, which matters for attention as much as for money. Third, the output tokens did not change: compaction is about what the model reads, not what it writes.

The third part of the script is the single most useful number in this chapter for a tool designer. One tool that returns 25,000 tokens at step 5 (a log dump, a full table, a whole document) and is then carried in the history for the remaining 26 steps adds 650,000 input tokens to the run, almost as much as the entire uncompacted run without it, and about $1.95 at the example price. The same tool trimmed to 800 tokens of relevant lines adds 20,800. That is the token-efficiency principle of Section 2.3 with a price on it, and it is why a response cap is a cost control.

Compaction and summarisation: how to do it without losing the plot

Compaction is the first lever and the one most likely to do quiet damage. The rules that keep it safe:

  1. Clear tool results first. Once a tool result has been read and acted on, the raw payload rarely matters again; the decision it led to is in the next model turn. Clearing old tool results is nearly lossless and often the only compaction you need.
  2. Keep the user's words verbatim as long as possible. Constraints and goals come from the user; a summary of them is a paraphrase by the model, and the paraphrase is where "never before 10:00" becomes "prefers mornings".
  3. Write the things that must survive into working memory before compacting. The scratchpad is re-inserted whole after compaction; the history is not.
  4. Give the summariser a recall-first prompt and test it on real traces. Plant facts, compact, ask. Only trim the prompt once nothing planted goes missing.
  5. Keep the last few turns verbatim. Recency bias is your friend here: the newest observation and the model's last thought are where the next decision comes from.
  6. Checkpoint before compacting. If the summary turns out to have dropped something, the checkpoint lets you rebuild from the full history.

Discussion

  1. How big should the budget be? Below the window by a wide margin, and set by measured accuracy, not by the vendor's maximum. My view: find the context length at which your per-step eval starts to slip, and set the budget below it.
  2. Compaction, notes or sub-agents? All three work; they fit different tasks. My view (following the post): compaction for conversational tasks, notes for milestone-driven work, sub-agents for parallel exploration; and tool-result clearing in every case.
  3. Should retrieved documents go at the start or the end? The U-curve says both ends are read well and the middle is not. My view: instructions at the start, retrieved material and the live plan near the end, just before the newest observation; and test the order on your eval, because models differ.
  4. Who pays for the tool definitions? They are paid on every turn by every run. My view: load per step when the catalogue is large, and treat definition tokens as a line item on the cost dashboard.

2.8 Putting the blocks together: the meeting example

One task, traced block by block. The request: "Set up a 45-minute design review with Priya and Tom next week. I prefer mornings but never before 10:00. Send the invitations."

The meeting-scheduling run, block by blockinstructionsmemorymodeltoolsstate1steprole: scheduler;rule: not before 10recall: user prefersmorningsplan: slots, room,book, invite(none yet)step 1/8, plan set,budget 82steptool guidance: findslots before bookingnote: candidatesMon/Wed/Thu/Fri 10picks find_free_slotsearliest_hour=10find_free_slots ->14 slotspending call,then cleared3stepoutput: event idin the summarynote: room 4Bpicks find_rooms,then create_eventcreate_event ->EVT-1187 Mon 10:00step 3/8, event idstored4stepconstraint: onlycompany addressesstore: Tom outon Mondayreads the bounce;decides to movesend_email ->ok + auto-replyerrors: 1, statusneeds_retry5stepstop: when theinvite is sentrecall: the 10:00rule (compacted!)picks update_eventWed 10:00update_event ->moved, invites sentstep 5/8, done;checkpoint finalevery step touches all five blocks; the trace in Chapter 8 records exactly these columns
The meeting-scheduling run traced block by block over five steps: what the instructions contributed, what memory recalled or stored, what the model decided, which tool ran and what the state recorded. At step 4 the invitation to Tom bounces; at step 5 the agent recovers, and the 10:00 rule, by then compacted out of the history, is recovered from long-term memory.
StepInstructionsMemoryModelToolsState
1role: scheduler; constraint: not before 10:00; stop: when invitations are sentrecalls the user's stored preference for mornings; writes the goal to the scratchpad; stores the 10:00 rulereads the request, writes a plan: find slots, check a room, book, invitenone yetstep 1 of 8, plan set, budget 8 steps
2tool guidance: find slots before bookingnotes the candidate slots at or after 10:00calls calendar_find_free_slots with earliest_hour=10returns 14 shared slots in a bounded listpending call recorded, then cleared with the result
3output format: event id in the summarynotes room 4Bcalls find_rooms, then calendar_create_event with an idempotency keyevent EVT-1187 created for Monday 10:00step 3 of 8; event id stored in state
4constraint: only company addressesstores "Tom out on Monday" as a factreads the tool result: invitations sent, but an auto-reply says Tom is out Monday; decides to moveemail_send returns ok plus the bounce text in the result (not a silent success)errors: 1; status: needs retry
5stop: when the invitations are sentrecalls the 10:00 rule from the store (the user's sentence was compacted at step 4)chooses Wednesday 10:00 from the remaining slots; calls update_eventevent moved, updated invitations sentstep 5 of 8; done; final checkpoint

Every step used all five blocks, and the two moments where the run could have gone wrong were both block failures averted by design. At step 4 a tool that returned "sent" without the bounce text would have ended the run with Tom absent: a tools failure (success on failure). At step 5 an agent that relied on the history alone would have had no 10:00 rule, because compaction had summarised the user's sentence away: a memory failure. The code in ch2_memory.py shows the second case literally.

What would break

If this block were weak...Symptom in the runWhere the trace shows itThe fix lives in
Missing memory (no store, compaction only)books Tuesday 09:00 at step 5 with a clear consciencethe context at step 5 has no 10:00 rulewrite constraints to the long-term store at step 1; re-insert the scratchpad after compaction
Bad tool description (do_calendar: "Calendar stuff.")calls the wrong calendar tool, or the right one with a free-text datetool-selection error; argument validation failurenamespaced names, when-to-use descriptions, strict schemas
Huge tool return (find_free_slots returns every calendar entry)step 2 observation of twenty thousand tokens; the plan from step 1 is "lost in the middle" by step 4observation size spike; later contradictionsa cap and filters in the tool; tool-result clearing
Vague instruction (no stop condition)after the invitations are sent, the agent "improves" the event, adds an agenda, emails againsteps continue after the goal; budget exhaustedan explicit done condition in the prompt and the loop
No state (free loop, no checkpoint)the process restarts after step 3 and books a second event, then sends duplicate invitationstwo event ids in the side-effect log; no record of the pending callcheckpoint after every step; record pending calls; idempotency keys
Weak model (routine model on the hard step)at step 4 it reads the bounce as success, or retries the same emailthe reasoning line ignores the auto-reply textroute recovery steps to the large model; make the tool put the error first in its result

Notice what the table does not say: "use a bigger model" appears once, as the last row, and only after five rows that a bigger model would not have fixed.

Exercises

  1. Your agent has 40 tools from four MCP servers, and tool-selection accuracy on your eval fell from 94% to 81% when the fourth server was added. List three remedies (namespacing, per-step tool loading, consolidation), the pros and cons of each, and the measurement that would tell you which one worked.
  2. Design the send_email tool for the meeting agent as a full definition: name, description, schema, return shape, error messages, idempotency and annotations. Then write the three eval cases you would use to test it, including one where the right answer is not to call it.
  3. A team proposes a router: a small model for every step, escalating to a large one when the small model's own confidence is below a threshold. Argue both sides: what does this save, what can go silently wrong, and what per-step eval would you insist on before shipping it?
  4. Your compaction prompt summarises the history every 15 steps. Plant five facts at steps 2, 5, 8, 11 and 14 and design the questions you would ask at step 30 to measure recall. Which facts would you expect to survive, which would you move to the long-term store or the scratchpad, and why?
  5. Re-run the thirty-step arithmetic of Section 2.7 for your own agent: your instruction length, your tool count, your observed observation size and your provider's prices. At what step does the uncompacted context cross your practical budget, and what is the cost of the single largest tool result in your traces re-sent for the rest of the run?

Key takeaways

  • An agent is assembled from five blocks: the model reasons, the tools act, the instructions shape, the memory remembers and the state says where the loop is. Each has its own failure modes and its own eval; when a run fails, find the block before you change anything.
  • Choose the model by measurement on your own tasks: tool-call accuracy, argument accuracy, refusal rate, latency p50 and p95, cost per task, parse-failure rate. Public leaderboards shortlist; your labelled set decides. Even the best models miss a quarter of the public benchmark's cases.
  • A router (small model for routine steps, large for hard ones) can cut cost and latency, at the price of two models to evaluate and a silent failure mode when the route is wrong. Build the per-step eval set first.
  • A tool is a name, a description and a JSON schema; the model emits a structured call, your code runs it and returns the result as a message. Design tools as interfaces: namespaced names, when-to-use descriptions, unambiguous arguments, bounded returns, errors that say what to do, idempotent writes, declared side effects. Consolidate endpoint wrappers into task-shaped tools.
  • MCP standardises how tools, resources and prompts are exposed to a host; it buys reach and costs control. Loading every tool definition into every turn does not scale; load what the step needs.
  • The system prompt is architecture: role and goal, hard constraints, tool guidance, output format, stop conditions and a few canonical examples, written at the right altitude, versioned and tested as code.
  • Memory lives in three places (the context, a scratchpad, a store) and holds three kinds of thing (episodic, semantic, procedural). Generative Agents gave the retrieval formula (recency, importance, relevance); MemGPT gave the paging analogy; production systems use compaction, structured notes and sub-agents.
  • State is control, not knowledge: step, plan progress, pending call, budget, errors, status. Checkpoint it after every step, record pending calls, and make writes idempotent, or a retry will do the task twice.
  • The model sees only the context. Budget it: fixed parts are paid every turn, the history grows every step, and attention fades in the middle long before the window is full. In the thirty-step example, compaction every ten steps cut input tokens by 44% and cost by 40%; one uncapped 25,000-token tool result cost almost as much as the whole run.
  • In the meeting example, the two near-failures were a tool that could have hidden an error and a memory that could have lost a constraint. "Use a bigger model" fixed neither.

References

Papers

  1. Shishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. Gonzalez. Gorilla: Large Language Model Connected with Massive APIs. NeurIPS 2024 (arXiv May 2023).
  2. Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, Thomas Scialom. Toolformer: Language Models Can Teach Themselves to Use Tools. NeurIPS 2023 (arXiv February 2023).
  3. Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, Maosong Sun. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. ICLR 2024 (arXiv July 2023).
  4. Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein. Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023 (arXiv April 2023).
  5. Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez. MemGPT: Towards LLMs as Operating Systems. arXiv October 2023.
  6. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the ACL, volume 12, 2024 (arXiv July 2023).

Engineering blogs and docs

  1. Fanjia Yan, Huanzhi Mao, Charlie Cheng-Jie Ji, Tianjun Zhang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez. Berkeley Function-Calling Leaderboard. Gorilla blog, UC Berkeley, last updated 19 August 2024; and the live leaderboard (V4, updated 12 April 2026 at the time of writing).
  2. Anthropic (Ken Aizawa and colleagues). Writing effective tools for agents, with agents. Engineering blog, 11 September 2025.
  3. Anthropic. Effective context engineering for AI agents. Engineering blog, 29 September 2025.
  4. Anthropic. Code execution with MCP: Building more efficient agents. Engineering blog, 4 November 2025.
  5. Anthropic. Building effective agents, Appendix 2, "Prompt engineering your tools". Research blog, 19 December 2024.
  6. OpenAI. A practical guide to building agents, "Configuring instructions", page 11. PDF guide, April 2025.
  7. Model Context Protocol. Architecture overview. modelcontextprotocol.io documentation, protocol version 2026-07-28.
  8. LangChain. Persistence. LangGraph documentation, accessed October 2026.

Code for this chapter

  1. code/agents/ch2_tools.py: one function-calling round trip with a real JSON schema and a stubbed model, and the tool-selection comparison on ten labelled requests; outputs in results/ch2_tools_stdout.txt and results/ch2_tools_select_stdout.txt.
  2. code/agents/ch2_memory.py: the fifty-line memory with a scratchpad, a summariser stub and a key-value store, with context sizes per step; results in results/ch2_memory.json.
  3. code/agents/ch2_math.py: the context budget, the thirty-step run with and without compaction, and the cost of one oversized tool result; results in results/ch2_math.json.
  4. code/agents/figs_ch2.py and code/agents/shots_ch2.py: the figures and the paper excerpts.

Next

→ Chapter 3: Single-agent reasoning patterns

With the blocks in hand, Chapter 3 asks how a single agent should think between tool calls. It reads ReAct closely (the interleaving of thought, action and observation that Chapter 1 introduced), then plan-and-execute (write the whole plan first, then carry it out), reflection and self-critique (Reflexion's verbal memory of failures), and tree search over possible next steps, and for each pattern it asks the questions this chapter asked of the blocks: what it costs in tokens and steps, when it is the right choice, what the papers measured, and how to tell from a trace that it is working.