RAG: From First Principles · Appendix

Appendix B · Every Metric on One Sheet

Notation: for one question, R = set of relevant chunk ids (from the golden set's sources, mapped to chunk ids), L = [l₁, l₂, …] = retrieved chunk ids in rank order, L@k = first k of them, rel(i) = 1 if lᵢ ∈ R else 0. All…

For each metric: what it measures, the formula, how to compute it (function names refer to code/ch10/metrics.py), what "good" looks like on a corpus like the handbook, and what to change when it is low. Chapter 10 has the code and worked examples; this is the reference you keep open.

Notation: for one question, R = set of relevant chunk ids (from the golden set's sources, mapped to chunk ids), L = [l₁, l₂, …] = retrieved chunk ids in rank order, L@k = first k of them, rel(i) = 1 if lᵢ ∈ R else 0. All metrics are computed per question and averaged over the golden set; report the mean and, ideally, the per-tag breakdown (numeric / multi_hop / unanswerable …).


B.1 Retrieval metrics (no LLM needed: run these first, they are free)

MetricMeasuresFormulaCodeGood on handbookIf low →
Hit rate@k (a.k.a. success@k)Did any relevant chunk make the top k?`1 ifL@k ∩ R> 0 else 0`hit_rate_at_k(retrieved, relevant, k)
Recall@kFraction of relevant chunks retrieved`L@k ∩ R/R
Precision@kFraction of retrieved that are relevant`L@k ∩ R/ k`precision_at_k
F1@kHarmonic mean of P@k and R@k2PR/(P+R)f1_at_k≥ 0.5Whichever of P/R is lower.
MRR (mean reciprocal rank)How high is the first relevant chunk?1 / rank_of_first_relevant (0 if none)mrr≥ 0.7Reranking is the direct fix; query rewriting; hybrid. MRR up while recall flat = ranking fixed, retrieval unchanged.
MAP (mean average precision)Precision averaged at each relevant hit`(1/R) Σ_{i: rel(i)=1} P@i`mean_average_precision
nDCG@kRanking quality with position discount (supports graded relevance)DCG@k = Σᵢ₌₁ᵏ rel(i)/log₂(i+1); nDCG = DCG / IDCG (IDCG = DCG of the ideal ordering)ndcg_at_k≥ 0.7Reranker; chunk ordering; if graded labels exist, tune to them. The standard metric in IR papers: know how to derive it.
Recall vs brute force (ANN recall)Is the index losing results the exact search would find?`ANN top-k ∩ exact top-k/ k` on a samplecompute with client.query_points(..., search_params=SearchParams(exact=True)) vs default

Reading them together

  • Hit rate high, recall low → finding some evidence but not all; multi-source questions suffer. Multi-query / decomposition.
  • Recall high, precision low → context is noisy; the generator pays in tokens and distraction. Rerank, lower k.
  • Recall high, MRR low → right chunks, wrong order. Rerank.
  • Everything low for numeric/code tags only → embeddings blur numbers; add BM25.
  • Everything low for one document only → chunking or parsing problem in that document.

B.2 Context quality (LLM-judged, per retrieved chunk)

MetricMeasuresFormulaCodeGoodIf low →
Context precisionOf the retrieved chunks, how many are actually useful for the answer: weighted by rankjudge each chunk useful ∈ {0,1} given question + reference answer; Σᵢ (P@i · usefulᵢ) / Σ usefulᵢ (RAGAS form) or simply Σ usefulᵢ / kcontext_precision_at_k (rank-weighted, doc-level labels; Chapter 10)≥ 0.6Reranker; smaller k; better chunk boundaries.
Context recallDoes the retrieved context cover the reference answer?split reference answer into statements; supported statements / total statementscontext_recall_docs (doc-level; the claim-level LLM version is faithfulness-style judging in llm_judges.py)≥ 0.85Retrieval recall levers; larger chunks or parent-document retrieval so the whole fact is present.
Context utilisation (RAGChecker)Of the relevant context that was retrieved, how much did the answer use?claims in answer supported by context / relevant claims available in context-≥ 0.7Generation problem: prompt ordering (lost in the middle), too much context, model too small.

B.3 Generation metrics (LLM-judged, on the answer)

MetricMeasuresFormulaCodeGoodIf low →
Faithfulness (groundedness)Are the answer's claims supported by the retrieved context?split answer into atomic claims; judge each supported ∈ {0,1} against the context only; supported / totalfaithfulness(answer, context)≥ 0.9; 1.0 on numericHallucination. Grounding prompt ("only from context"); require citations by chunk id and validate them; faithfulness gate that rewrites or refuses; lower verbosity; check context actually contained the fact (if not, this is a retrieval failure disguised).
Answer relevanceDoes the answer address the question (regardless of truth)?generate n questions from the answer, mean cosine(generated, original) (RAGAS); or direct 1–5 judge with rubricanswer_relevance≥ 0.8Prompt: answer the question first; stop the model from summarising the whole context; check query rewriting did not change intent.
Answer correctnessDoes the answer agree with the reference?judge correct ∈ {0, 0.5, 1} with the reference; plus cheap checks: keywords all present (keyword_hit), numeric exact matchanswer_correctness(question, answer, reference) + keyword_hit(answer, keywords)≥ 0.8If faithfulness is high but correctness low → wrong/insufficient context (retrieval) or conflicting sources (add recency metadata). If faithfulness low too → generation.
CompletenessDid it cover all parts of a multi-part reference?reference statements covered / totalpart of answer_correctness≥ 0.8Decomposition; larger k for multi_hop; ask for all parts explicitly.
Citation accuracyDo cited chunk ids actually support the sentence that cites them?judge per citation; valid / total citationscitation_accuracy (code/ch10/llm_judges.py)≥ 0.9Structured output with citations; validate ids exist; penalise uncited claims in the prompt.
Refusal precision / recallRefuses when it should, answers when it canover golden answerable flag: refusal-precision = correct refusals / all refusals; refusal-recall = correct refusals / unanswerable questionsis_abstention + the answerable flag in run_eval.py (Chapter 10/12)both ≥ 0.8Low recall (answers made-up stuff): retrieval score threshold, "answerable" field, grounding prompt. Low precision (over-refuses): threshold too strict, prompt too cautious, retrieval missing real answers.
Conciseness / lengthTokens in the answerlen(tokens)trivial30–150 tokens for factoidPrompt; lower verbosity settings. Long answers also inflate judge scores (verbosity bias).

B.4 End-to-end and business metrics

MetricMeasuresHowGoodIf low →
Exact / keyword matchcheap correctness proxyall keywords present (case-insensitive, normalised numbers)≥ 0.85Same as correctness; check normalisation (₹1,500 vs 1500).
Task success ratedid the user get what they neededthumbs up/down, "was this helpful", follow-up-question rate> 70% helpfulLook at the traces of the failures; cluster them; fix the biggest cluster.
Escalation / deflection ratesupport use cases: tickets avoidedtickets before vs afterdomain-specificCoverage analysis: which questions have no answer in the corpus (Chapter 12).
Retrieval-call rate (agents)fraction of factual questions where the agent actually retrievedfrom traces~100% on factualForce retrieval structurally (Chapter 16).
Steps / tokens per task (agents)loop costfrom traces / usage_metadata≤ 3 steps medianLimits, termination condition, better tools.

B.5 Operational metrics

MetricMeasuresHowTypical targetIf bad →
Latency p50 / p95 / p99end to end and per stage (embed, retrieve, rerank, generate)timers in the trace; time.perf_counter() per stagep95 < 2–3 s for chat; retrieval < 100 msStage that dominates: generation → stream, smaller model, shorter context; rerank → fewer candidates, lighter model, GPU batching; retrieval → ef, quantisation, payload indexes, caching.
Time to first tokenperceived latencystream and time the first chunk< 1 sStream; move checks after streaming; prompt caching.
Cost per query$(input_tokens × p_in + output_tokens × p_out) from usage_metadata + embedding + infra/QPSknow your numberShorter context (k, chunk size), smaller model for easy questions, caching, prompt-prefix ordering for cache hits.
Tokens per query (input/output)the driver of cost and latencyusage_metadatae.g. 1,500 in / 200 outSame.
Throughput (QPS) and error ratecapacity, reliabilityload test; 429/5xx countserror < 0.5%Async I/O, connection pools, provider fallback, retries with backoff.
Index freshness lagtime from source change to searchabletimestamp diff< 5 minEvent-driven ingestion; smaller batches.
Cache hit ratesemantic/prefix cache effectivenesshits / lookups20–40% semantic; > 70% prefixStable prompt prefix first; threshold tuning; normalise queries.

B.6 Evaluation-process metrics (is your judge trustworthy?)

MetricMeasuresFormulaTargetIf low →
Judge–human agreementraw agreementmatches / n on a human-labelled sample (50–100 items)> 85%Rubric with explicit criteria; claim-level decomposition; stronger judge model; different family from the generator.
Cohen's kappaagreement corrected for chanceκ = (p_o − p_e)/(1 − p_e) where p_o = observed agreement, p_e = expected by chance from marginalsκ > 0.6 (substantial), > 0.8 (near-perfect)Same as above; binary labels are easier to agree on than 1–5 scales.
Position-swap consistencypairwise judge stabilityfraction of pairs where verdict is the same after swapping A/B> 90%Always evaluate both orders and average; use absolute rubric scoring instead of pairwise.
Judge variancerun-to-run noisestd of scores over repeated runssmall relative to the effect you care aboutFixed prompts, lower randomness, more items.
Golden-set coveragedoes the eval set represent trafficfraction of real-query clusters with ≥ n golden itemsall major clustersAdd real questions from logs; regenerate synthetic ones per new document.

B.7 The formulas you should be able to write on a whiteboard

plain text
precision@k = |retrieved@k ∩ relevant| / k
recall@k    = |retrieved@k ∩ relevant| / |relevant|
MRR         = mean over questions of 1 / rank(first relevant)      (0 if none)
AP          = (1/|relevant|) · Σ_{i : rel(i)=1} precision@i ;  MAP = mean(AP)
DCG@k       = Σ_{i=1..k} rel(i) / log2(i + 1)
nDCG@k      = DCG@k / IDCG@k
faithfulness = supported claims / total claims           (claims from the answer, judged vs context)
context recall = reference statements supported by context / total reference statements
RRF(d)      = Σ_lists 1 / (60 + rank_list(d))
BM25(q,d)   = Σ_t IDF(t) · tf·(k1+1) / (tf + k1·(1 − b + b·|d|/avgdl)),  k1≈1.5, b≈0.75
cosine(a,b) = a·b / (|a||b|)
kappa       = (p_o − p_e) / (1 − p_e)

B.8 A minimal eval report template

plain text
Config: collection=handbook_v3  chunk=800/120  embed=text-embedding-3-small  k=5  rerank=none  model=gpt-5.4-mini
Golden: 47 questions (42 answerable, 5 unanswerable)   Judge: gpt-5.4-mini, rubric v2

Retrieval        hit@5 0.93   recall@5 0.86   precision@5 0.41   MRR 0.78   nDCG@5 0.81
Generation       faithfulness 0.94   relevance 0.88   correctness 0.83   keyword-hit 0.86
Refusal          precision 0.83   recall 1.00
Cost/latency     1,420 in / 96 out tokens   $0.0015/q   p50 1.9 s   p95 3.4 s
Worst tags       multi_hop: recall 0.62   numeric: correctness 0.75
Judge check      agreement 0.90 (n=40)   kappa 0.78

Write one of these per change. The diff between two reports is the result of your experiment; anything not in the report did not happen.

B.9 Before you trust any number above

Every metric in this appendix is an estimate from a sample, computed against labels you invented, and half of them are produced by a stochastic judge. Four checks separate a result from a story: all are derived and measured in Chapter 21 and Chapter 26.

CheckQuestion it answersWhere
Bootstrap CI on every headline metricis this difference bigger than the noise? On 42 questions the 95% interval on recall is about ±0.08: a 2-point "win" is nothingCh 21 §21.5, code/ch21/significance.py
Paired permutation test, not two independent meansdid config B beat config A on the same questions? Pairing removes question difficulty as a confoundCh 21 §21.5
Label granularity stated out louddocument-level labels score a retrieval as perfect when the answer-bearing chunk was never returned; on this corpus they inflate recall by ~5 pointsCh 21 §21.3, code/ch21/chunk_level_labels.py
Judge stability re-measured as context growsthe same judge scoring the same objectively-correct citation was 100% self-consistent on a 1-chunk context and 60% on a 10-chunk one: judge error grows with exactly the variable you vary when tuning kCh 26 §26.5, code/ch26/judge_bias.py

The practical rule: no single metric moving alone is a result. Trust changes where several related metrics move together, where the CI excludes zero, and where you can name the mechanism.

And the sentence to have ready in an interview: recall@k is stage one of four: retrieved, then survived context assembly, then used by the model, then correct. Quoting retrieval recall as an accuracy figure claims the other three stages are lossless.