RAG: From First Principles
To terabyte-scale agentic systems
Retrieval-augmented generation built and measured step by step. Embeddings, chunking, hybrid search, reranking, evaluation, hallucination, agentic RAG with LangGraph, and ingesting a million PDFs.
A complete, self-contained course on Retrieval-Augmented Generation, written to be read start to finish, in order, with code you run and numbers you measure at every step. One learning path, one corpus, one stack.
Stack (current as of September 2026): Python 3.12 · LangChain 1.4 · LangGraph 1.2 · OpenAI (chat and embeddings) · Qdrant 1.19 · FastAPI. Chapters 1 to 13 use LangChain only; LangGraph enters in Chapter 14, FastAPI and Docker in Chapter 17. Every API was checked against the official docs, so the old RetrievalQA / create_react_agent / MemorySaver world is not used anywhere (see Appendix D for the old to new table).
What this is: the curriculum for someone who has to build and defend a RAG system. Every chapter ends with the questions engineers are actually asked about it, with answers.
What this is not: an API tour. You will learn why each piece exists, how to measure whether it works, and what to change when it does not.
The one idea that organizes everything
A RAG system is only as good as the worst of three stages, and you cannot fix a stage you have not measured:
| Stage | Question it answers | Measured by | Chapters |
|---|---|---|---|
| Ingestion | Did the right facts end up as retrievable chunks? | chunk stats, coverage | 5, 6, 7, 12 |
| Retrieval | Given a question, did the right chunks come back, near the top? | recall@k, precision@k, MRR, nDCG | 4, 8, 9, 10 |
| Generation | Did the model answer from those chunks, correctly, and admit when it can't? | faithfulness, correctness, refusal accuracy | 3, 10, 11, 12 |
Everything in this book is either how to measure one of these, or what to do about it. Agentic RAG (Part 5) is what you build when a single pass through the three stages is not enough and the system has to decide to retrieve again, differently, or not at all.
Parts I to VII build the system. Part VIII (Chapters 21 to 30) is the layer underneath: HNSW built from scratch, rank fusion derived, rerankers measured, and hallucination causes injected on purpose, so you can explain why each component works and not only how to wire it up.
Before you start
- Prerequisites. Python at the level of "I can write a function, a class, and read a stack trace." No ML background needed; Chapter 1 gives you the vocabulary.
- Setup. Five minutes:
uv sync, put your OpenAI key in.env, run one script. Qdrant runs embedded (no Docker needed) until Chapter 17. - Cost. The whole book, run once end to end, costs roughly $1 to $3 in OpenAI usage with the default models. Evaluation loops are the expensive part; they have
--limitflags. - Time. About 40 to 60 hours including exercises. The study plan has a 4-week plan and a 5-day sprint.
- Corpus. Every chapter uses the same data: the internal handbook of a fictional robotics company (14 documents) and a golden set of 47 questions with reference answers, including 5 the corpus cannot answer and one deliberate contradiction. See the corpus.
How to read this book
- In order. Each chapter's code assumes the previous chapter's ideas, and the corpus, golden set and helpers are introduced progressively.
- Run everything. Each chapter has a "Run it" section with the exact command and the output you should see. The numbers you measure are the course.
- Keep a
notes.md. For every evaluation you run, record the config, the numbers, and why you think they moved.
Chapters
- Setup
Five minutes. You need Python 3.12+ via uv, and an OpenAI API key.
- Study Plan
Three versions. Pick one and put it in your calendar.
- The Corpus and the Golden Set
Fourteen markdown documents making up the internal handbook of Lumora Robotics, a fictional company that builds warehouse robots (Atlas), fleet software (Beacon) and analytics (Compass). About 5,200 words. Written so…
- Chapter 1 · Why RAG Exists
A large language model (LLM) answering a question draws on two different sources of information, and it helps to keep them apart from the first page:
- Chapter 2 · Embeddings: Meaning as Geometry
Code for this chapter: code/ch02/embeddingsplayground.py.
- Chapter 3 · Your First RAG (LangChain Only)
Code: code/ch03/firstrag.py.
- Chapter 4 · Qdrant: A Real Vector Database
Code: code/ch04/qdrantrawclient.py, code/ch04/qdrantlangchain.py.
- Chapter 5 · Chunking, Measured
Code: code/ch05/chunkinglab.py, code/ch05/semanticchunker.py.
- Chapter 6 · A Generic, Reusable RAG Pipeline
Code: code/ch06/ingest.py, code/ch06/retrieve.py, code/ch06/generate.py, code/ch06/ragcli.py. The finished pipeline is also packaged as ragbook/index.py, which every later chapter calls.
- Chapter 7 · More Data: Growing and Maintaining the Corpus
Code: code/ch07/downloadgutenberg.py, code/ch07/synthesizedocs.py, code/ch07/ingestbatch.py.
- Chapter 8 · BM25 and Hybrid Search
Chapters 2–7 built a dense retriever: every chunk becomes one vector, the query becomes one vector, and we return the nearest ones by cosine similarity. It is excellent at meaning: "how fast can the robot move" finds the…
- Chapter 9 · Advanced Retrieval: Re-ranking, Query Transforms, Parent/Child, Metadata
Chapter 6's retriever was one step: embed → nearest-k. Every production system ends up as a funnel:
- Chapter 10 · Evaluating RAG: Every Metric, How to Compute It, How to Move It
Every chapter so far ended with "measure it". Here is why: a RAG system has at least eight knobs (chunk size, overlap, k, embedding model, hybrid on/off, re-ranker, prompt, LLM) and every knob interacts with every other…
- Chapter 11 · Stopping Hallucination
The pitch for RAG is "the model answers from your documents, so it cannot make things up." That is half true. Retrieval fixes the knowledge problem: the model no longer has to remember your PTO policy. It does not fix…
- Chapter 12 · When the Data Isn't There
Every RAG demo is built on questions the corpus can answer. Every RAG product gets questions it cannot: the policy that was never written down, the product that does not exist, the thing that lives in a different system…
- Chapter 13 · Observability: Logging, Tracing, Cost and Latency
"Why did it answer that?" A RAG answer depends on a query embedding, a nearest-neighbour search, a prompt template, a stochastic model and whatever was in the index at that second. If you did not record those, you cannot…
- Chapter 14 · LangGraph Basics: RAG as a State Machine
Everything before this chapter used LangChain alone: a retriever, a prompt, a model, glued together with Python. That is enough for "2-step RAG": retrieve, then generate. It stops being enough the moment you want the…
- Chapter 15 · Agentic RAG
The interviewer asked: "What is the difference between agentic RAG and generative RAG?" The candidate said "agentic can reason and decide how to retrieve." True, but here is the version that shows you have built one:
- Chapter 16 · Multi-Agent Systems and Guard Rails
Why it happens. Nobody owns termination. Each agent's job is phrased as "find what is missing": the planner always proposes one more search, the analyst always finds one more gap, and the graph has a cycle with no exit…
- Chapter 17 · Production: FastAPI, Async, Docker
New libraries in this chapter: FastAPI (web framework), uvicorn (the ASGI server), httpx (async HTTP client), Docker. Everything RAG-related is unchanged from Chapter 6.
- Chapter 18 · Terabyte Scale
Assumptions (state them, then change them): ~4 characters per token, 800-character chunks with 120 overlap (what ragbook.index uses), text-embedding-3-small at 1536 dims and $0.02 per 1M tokens, a sustained embedding…
- Chapter 19 · What's New in RAG, and Is RAG Dead?
This chapter is prose with pointers. Written September 2026; facts are cited inline. Where something is a vendor claim rather than an independent result, it says so.
- Chapter 20 · The Interview Question Bank
Answer shape that works in interviews: definition in one sentence → how it works → the trade-off → a number or name → what you would do in production. Most answers below follow it.
- Chapter 21 · Retrieval Metrics Under the Microscope
No. It means that for 95% of the questions in that evaluation set, under that definition of relevance, at least the labelled evidence appeared somewhere in the top 5 results. It is a claim about one stage of a four-stage…
- Chapter 22 · Vector Search Internals: HNSW, IVF and Quantization
Chapter 4 gave you HNSW in one page and a table of three knobs: enough to configure a collection. That is the API layer. This chapter is underneath it: we build the graph, break it deliberately, and measure what each…
- Chapter 23 · Fusion: Score Distributions, RRF, and Learned Sparse Retrieval
Chapter 8 built BM25, showed the RRF formula, and wired up hybrid search three ways. It asserted that "dense scores are cosines in ~[0.2, 0.8]; BM25 scores are unbounded: you cannot add them." This chapter proves it…
- Chapter 24 · Rerankers: Bi-encoders, Cross-encoders and Late Interaction
Here is the whole argument in one sentence: a bi-encoder must decide what a chunk means before it has seen your question.
- Chapter 25 · Choosing and Evaluating an Embedding Model
The interview question is "how do you choose an embedding model?" The weak answer lists properties: cost, dimensions, multimodality, whether it is open source. Those are inputs. The strong answer is a procedure that ends…
- Chapter 26 · Generation Evaluation and Hallucination Forensics
The question is: the retrieved chunks are correct, but the model still hallucinates: what do you do? The tempting answer is "lower the temperature".
- Chapter 27 · Retrieval Debugging: the Rank-17 Playbook
The instinct is understandable. The document is at rank 17, k is 5, so make k 20 and the bug is gone. Here is what you actually bought:
- Chapter 28 · Documents in the Real World: PDFs, Tables, Diagrams
Chapter 5 gave this one section and the honest summary "parsing is where the battle is". This chapter fights the battle.
- Chapter 29 · Context Construction: What Actually Reaches the Model
A million-token context window is not a licence to use it. Four reasons, in order of how often they bite:
- Chapter 30 · Ingesting a Million PDFs
Chapter 18 sized the index: vectors, RAM, quantization. This chapter is about getting the documents into it, which is where the actual engineering time goes.
- Appendix A · The Research Papers, in the Order to Learn Them
Time budget if you read only this appendix: about 3 hours. The two-week plan and the "only 10" list are at the end.
- Appendix B · Every Metric on One Sheet
Notation: for one question, R = set of relevant chunk ids (from the golden set's sources, mapped to chunk ids), L = [l₁, l₂, …] = retrieved chunk ids in rank order, L@k = first k of them, rel(i) = 1 if lᵢ ∈ R else 0. All…
- Appendix C · Glossary
One or two lines per term. Chapter numbers point to where the term is used in earnest.
- Appendix D · API Cheatsheet (LangChain 1.x · LangGraph 1.x · langchain-qdrant · qdrant-client)
Keep this open while coding. Section D.1 is the 90% you use every day.