RAG: From First Principles · Part 1 · Foundations

Chapter 1 · Why RAG Exists

A large language model (LLM) answering a question draws on two different sources of information, and it helps to keep them apart from the first page:

Goal: by the end of this chapter you can explain, without notes, what problem retrieval-augmented generation solves, why the alternatives (fine-tuning, giant context windows) do not make it obsolete, and what every box in a RAG pipeline is for. You will also have the vocabulary the rest of this book uses. No code in this chapter: it is the only one.


1.1 Four limitations that motivate RAG

A large language model (LLM) answering a question draws on two different sources of information, and it helps to keep them apart from the first page:

plain text
                 LLM answering a question
                          │
               ┌──────────┴──────────┐
               │                     │
     parametric knowledge     in-context information
        (model weights)        (the current prompt)
               │                     │
        learned during        user text, retrieved docs,
           training           tool results, conversation

Parametric knowledge is what training wrote into the weights. In-context information is whatever you put in front of the model for this one call. A model can know something in its weights and not have it in the context, or know nothing about a topic from training and still answer correctly because you supplied the facts in the prompt. RAG works almost entirely through the second path. Relying on either path alone runs into four limitations.

1. Its parametric knowledge becomes stale. An LLM learns facts during training, but its weights do not update when the world changes. If your company released Beacon 4.2 yesterday, a model trained earlier cannot know those release notes from its parameters alone. It can still use them perfectly well if you put them in the prompt: the problem is not that the model cannot understand new facts, it is that new facts never get written into its weights. They have to be supplied at inference time, through the prompt, a tool, or retrieval.

2. You cannot rely on its weights to contain your private data. Your handbook, tickets, contracts, databases and newly written documents are generally not available to a model as authoritative knowledge. Even if something related appeared in its training data, you cannot assume it was memorised accurately, that it is the current version, or that the model can recall it reliably. Ask "how many PTO days do I get" and a model with no context can only answer for some company, not yours.

3. It can generate unsupported information. An LLM produces text by repeatedly choosing a likely next token given everything before it, P(xt∣x1,…,xt−1)P(x_t \mid x_1, \ldots, x_{t-1}). Nothing in that objective checks a claim against a source of truth. So when evidence is missing or ambiguous, a model may produce a fluent, plausible answer instead of reliably saying "I don't know". Models can abstain, and good prompting and training make them do it more often, but abstention is not guaranteed. This failure is called hallucination, and it is hard to spot because fluency and correctness are different properties: ask for a policy that does not exist and you may well get one, with numbers.

4. Its working context is finite and expensive. The context window limits how much text the model can process in one call (the prompt plus the answer). Even with million-token windows, every input token costs money and time, and models do not always use information equally well across a very long input. Liu et al. measured this in "Lost in the Middle" (2023): on multi-document question answering, accuracy was often highest when the relevant passage sat at the start or end of the context and dropped markedly when it sat in the middle. Newer models vary in how much they suffer from it, but pasting your entire wiki into every request is neither free nor reliable.

RAG helps with all four through one architectural idea: retrieve the relevant text at question time and put it in the model's context. Instead of requiring the model to remember an entire handbook in its weights, or placing the entire handbook in every prompt, the system retrieves the few passages most likely to contain the answer. The model no longer has to remember your handbook; it has to read the three paragraphs you hand it and answer from those. Lewis et al., who coined the term, framed exactly this split as parametric memory (the weights) plus non-parametric memory (a retrievable index) (Lewis et al., 2020).

Retrieval does not guarantee a correct answer. Every step can fail:

plain text
the right document exists
        ↓
the retriever finds it
        ↓
the right chunk reaches the prompt
        ↓
the model reads it correctly
        ↓
the model makes only claims the chunk supports
        ↓
correct answer

What RAG gives you is a better foundation: current external knowledge, a small targeted context, provenance for every answer, and an answer you can check against its sources. Most of this book is about measuring and hardening each of those arrows.

1.2 The three ways to give a model new knowledge

Fine-tuningLong context ("just paste everything")RAG
Howcontinue training the model on your documentsput all documents in every promptretrieve the relevant few pieces per question, put only those in the prompt
Updatesretrain (hours–days, $$)instantinstant: re-index the changed document
Cost per questionlow (short prompt)very high (whole corpus every time)low–medium (a few chunks)
Scales towhatever fits in training~1M tokens ≈ 750k words ≈ a couple of thousand pagesterabytes: retrieval is the filter
Provenance ("where did that come from?")none: it is in the weightspossible but the model must find itbuilt in: you know exactly which chunks were shown
Access controlimpossible per userpossible but wastefulnatural: filter retrieval by the user's permissions
Hallucinationstill happens; fine-tuning teaches style far better than factsreduced, if the model finds the factreduced, and checkable against the retrieved text
Good fortone, format, a new skill, a domain's jargonsmall, stable corpora; one-off analysis of a documentfacts that change, private data, large corpora, anything that needs citations

The honest summary: fine-tuning changes how a model behaves; RAG changes what it knows right now. They are not rivals: production systems often do both (fine-tune a small model to follow your answer format, RAG to feed it facts). Long context is a genuine competitor for small corpora and a genuine complement inside RAG (bigger context lets you retrieve more generously). Chapter 19 returns to the "is RAG dead?" debate with 2026 evidence; the short answer is that retrieval got more important as agents started reading more, not less.

1.3 The RAG loop

Two independent pipelines. The first runs when documents change; the second runs when a user asks something.

plain text
INGEST (write path - runs when data changes)

  documents ──► parse ──► chunk ──► embed ──► store
  (PDF, md,     (text +   (pieces   (vector    (vector DB:
   HTML, DB)     metadata) ~200–800  per chunk) vector + text
                           tokens)              + metadata)

QUERY (read path - runs per question)

  question ──► embed ──► search ──► (rerank) ──► top-k chunks ──► prompt ──► LLM ──► answer
              (same      (nearest   (optional:   (3–10 pieces    (system     (reads   (+ citations)
               model!)    vectors,   a stronger   of text)        rules +     the
                          + filters) model         	             context +   context,
                                     re-orders)                   question)  answers)

Every chapter of this book is about one box in this diagram, or about measuring whether a box is doing its job:

BoxWhat goes wrong thereChapter
parsetables destroyed, headers lost, PDFs read in the wrong order5, 7
chunkanswer split across two chunks; chunk too big to be specific5
embedwrong model, docs and queries embedded differently, cost2
storecan't filter by metadata, too slow at scale, no persistence4, 18
searchsemantic search misses exact terms (IDs, codes, names)8
rerank / top-kright chunk retrieved at rank 9, k was 59
promptmodel ignores the context, or answers from memory11
LLMhallucinates, refuses, invents citations11, 12
the whole loopyou feel it's better but can't prove it10, 13
the loop, but the model decides when to searchagents, loops, budgets14–16

1.4 Vocabulary

You will use these words in every interview about RAG. Learn the precise meaning.

  • Token: the unit models read and bill by. Roughly ¾ of an English word. Common words are one token; "Retrieval" is two or three (Ret/rie/val in OpenAI's cl100k_base tokenizer) and "Lumora" is two or three.
  • Embedding: a fixed-length list of numbers (a vector) produced by an embedding model from a piece of text, arranged so that texts with similar meaning get vectors that point in similar directions. text-embedding-3-small gives 1,536 numbers per text.
  • Cosine similarity: the cosine of the angle between two vectors. 1 = same direction, 0 = unrelated, −1 = opposite. For normalized vectors (length 1) it equals the dot product. This is the number that ranks search results. Chapter 2 computes it by hand.
  • Chunk: a piece of a document small enough to be one search result and one prompt ingredient. Typically 200–800 tokens. Chunking strategy is the single highest-leverage decision in a RAG system (Chapter 5).
  • Vector database: a store that keeps vectors together with their text and metadata and answers "which stored vectors are closest to this one?" fast, even for millions of vectors. We use Qdrant.
  • ANN / HNSW: approximate nearest neighbour search. Comparing a query against every stored vector is exact but O(n). HNSW (Hierarchical Navigable Small World) is a graph index that finds almost the same neighbours in roughly O(log n). "Approximate" is the price of speed; you tune how approximate (Chapter 4, 18).
  • Top-k: the k best-scoring chunks the retriever returns. k is a dial: too small and you miss the answer; too large and you dilute the context and pay for tokens.
  • Retriever: the component that turns a question into a list of chunks. Dense (vector) search is one retriever; keyword search (BM25) is another; hybrid combines them (Chapter 8).
  • Reranker: a second, slower model that re-scores the top-k candidates with the question and each chunk together, producing a better order than embedding similarity alone (Chapter 9).
  • Context window: the maximum tokens an LLM can take in one call, prompt plus answer.
  • Grounding: making the model's answer rest on provided text. A grounded answer is one whose every claim can be traced to a retrieved chunk. Faithfulness is the metric for it (Chapter 10).
  • Hallucination: a fluent, confident, false statement. In RAG it splits into two: the retriever brought the wrong text (fix retrieval) and the model ignored or embellished the right text (fix the prompt/model). Diagnosing which is Chapter 11.
  • Golden set: a list of questions with known correct answers and known source documents. Without one you cannot measure anything (Chapter 10). Ours is data/golden/qa.yaml.
  • Agentic RAG: instead of "always retrieve once, then answer", the model decides whether to search, what to search for, whether the results are good enough, and whether to search again. Implemented with LangGraph in Part 5.

1.5 A transformer, in three sentences, and why it matters here

Interviewers ask this. Here is the answer at the right depth.

A transformer processes all tokens of its input at once and, in each layer, lets every token look at every other token through self-attention: a learned weighting of "which other tokens matter for understanding this one". Stacking these layers turns raw tokens into contextual representations, which is why "bank" next to "river" ends up represented differently from "bank" next to "loan". The 2017 paper Attention Is All You Need introduced it, and it displaced recurrent networks because attention is parallel (fast to train on large data) and has no distance penalty (a token at position 1 can attend to position 900 directly).

Why it matters for RAG: an embedding model is a transformer whose output for a text is pooled into one vector. Because the representation is contextual, "PTO", "annual leave" and "vacation days" land close together. That is the entire reason semantic search works: and also the reason it fails on things transformers do not treat as meaning, like a part number or a person's name, which is why Chapter 8 adds keyword search back.

1.6 Map of this book

  • Part 1: Foundations (Ch 1–3). Vocabulary, embeddings by hand, then the smallest complete RAG in ~70 lines of LangChain.
  • Part 2: Building (Ch 4–7). Qdrant, chunking (measured), a reusable pipeline, and growing the corpus.
  • Part 3: Retrieval (Ch 8–9). BM25, hybrid search, reranking, query rewriting, metadata filtering.
  • Part 4: Quality (Ch 10–13). Every metric with its formula and how to move it, hallucination, "not enough data", observability.
  • Part 5: Agentic (Ch 14–16). LangGraph, agentic RAG, multi-agent guardrails (loops, budgets, RAG-first enforcement).
  • Part 6: Scale (Ch 17–18). Serving, security, Docker, and the maths of a terabyte.
  • Part 7: Frontier (Ch 19–20). What changed in 2025–26, and an interview question bank.
  • Appendices. Papers in reading order, metrics cheat-sheet, glossary, API cheat-sheet.

Everything runs against one corpus: a fictional company's handbook in data/handbook/ (14 documents) and 47 golden questions in data/golden/qa.yaml. The corpus is small on purpose: you can read all of it in twenty minutes, which means you can check every answer the system gives. That habit is the course.

Run it

Nothing to run. Read data/handbook/02-pto-and-leave-policy.md and data/handbook/14-faq.md now and notice that they disagree about PTO days (24 vs 20). Remember that; it comes back in Chapters 3, 11 and 12.

bash
cat data/handbook/02-pto-and-leave-policy.md | head -20
grep -n "PTO" data/handbook/14-faq.md

Exercises

  1. Write down, from memory, the two pipelines of the RAG loop with every box. Compare with §1.3. Do it again tomorrow.
  2. Pick a question you might ask about your own employer. For each of the four LLM limits in §1.1, say whether it applies to that question.
  3. For each row of the comparison table in §1.2, think of one real situation where that column is the right choice. (Fine-tuning has one; find it.)
  4. Read the golden set data/golden/qa.yaml. Which questions do you predict will be hard for a naive system and why? Write the ids down; check your predictions in Chapter 10.

Interview questions

Q: Explain RAG end to end. Two pipelines. Ingestion: parse documents into text with metadata, split into chunks of a few hundred tokens, embed each chunk with an embedding model, store vector + text + metadata in a vector database. Query: embed the user's question with the same model, find the top-k nearest chunks (optionally filtered by metadata and reranked), build a prompt with a system instruction, the chunks as numbered context and the question, and let the LLM answer from that context with citations. Evaluate both halves separately: retrieval with recall@k, generation with faithfulness and correctness.

Q: Why embeddings? Why not just keyword search? Keyword search matches strings; embeddings match meaning, so "annual leave" finds the PTO policy. They fail in opposite places (embeddings are weak on exact identifiers, keywords on paraphrase) which is why production systems use both (hybrid search).

Q: Why store the embeddings? Can't we compute them at query time? The corpus side must be embedded once and searched many times; recomputing millions of vectors per query is impossibly slow and expensive. Storing them in a vector DB with an ANN index makes each query a millisecond graph walk instead of a full scan. Only the question is embedded at query time.

Q: RAG vs fine-tuning: when would you pick which? Fine-tune to change behaviour: tone, output format, a domain's jargon, a task the base model does poorly. Use RAG to change knowledge: facts that change, private data, anything that needs a citation or per-user access control. Fine-tuning does not reliably teach facts and cannot be updated in minutes; RAG cannot teach a new skill. Often you do both.

Q: What is a transformer, in a few sentences, and why does it matter for embeddings? A neural network that processes a whole sequence in parallel and uses self-attention so every token's representation is shaped by every other token; introduced in "Attention Is All You Need" (2017); replaced RNNs because it trains in parallel and handles long-range dependencies. Embedding models are transformers whose contextual token representations are pooled into one vector: that contextuality is why semantically similar texts end up nearby.

Q: Won't million-token context windows kill RAG? For a corpus that fits, long context is a real option and a simpler one. It does not fit for most enterprises (terabytes), it costs the whole corpus in tokens per question, it cannot enforce per-user access, and models measurably degrade at using facts buried in very long inputs. In practice long context made retrieval more useful: you can afford to retrieve more, and agents read more per step.

Q: What is hallucination in a RAG system and what are its two causes? A confident false statement. Either retrieval failed (the right text was never shown, so the model filled the gap) or generation failed (the right text was shown and the model ignored, misread or embellished it). You diagnose which by looking at the retrieved chunks first: that is why every script in this book prints them.

Key takeaways

  • An LLM answers from two sources: parametric knowledge in its weights and in-context information in the prompt. RAG works through the second.
  • RAG is motivated by four limitations: stale weights, private data the weights cannot be trusted to hold, unsupported (hallucinated) answers, and a finite, costly context. Retrieval helps with all four by putting the right text in the prompt; it does not guarantee a correct answer, because every step from document to answer can fail.
  • Fine-tuning changes behaviour; RAG changes knowledge. Long context is a complement, not a replacement: a million-token window is roughly two thousand pages, and you pay for all of them on every question.
  • Two pipelines: ingest (parse → chunk → embed → store) and query (embed → search → rerank → prompt → LLM). Every RAG bug lives in exactly one box.
  • The same embedding model must be used for documents and queries.
  • You cannot improve what you do not measure; a golden set comes before optimisation.

Next

→ Chapter 2: Embeddings: Meaning as Geometry

Chapter 2 makes "embedding" concrete: you will embed sentences, compute cosine similarity by hand, and watch nearest-neighbour search work (and fail) on real vectors.