Topic map
Attention and Transformers
- Self-attention
Every token looks at every other token and mixes in what it finds. The one idea everything else here builds on.
- Multi-head attention
Several attention heads run in parallel, each looking for a different kind of relation.
- MQA, GQA, MLA
Ways to share keys and values across heads so the KV cache gets smaller.
- Sliding-window and sparse attention
Let each token look only at a window or a pattern of others, to cut the quadratic cost.
- Linear attention and hybrids
Attention rewritten so its cost grows linearly, and models that mix it with the standard kind.
- The Transformer: encoder and decoder
The 2017 design: an encoder that reads and a decoder that writes, built only from attention and small MLPs.
- Position embeddings
Attention has no sense of order, so a learned vector per position is added to every token.
Pre-training (BERT)
- BERT
A Transformer encoder that reads text in both directions, pre-trained once and fine-tuned for any task.
- Masked language modelling
Hide 15% of the words and train the model to guess them from both sides.
- Pre-train, then fine-tune
Train once on a huge dataset, then adapt a copy to each task with one small new layer.
- The [CLS] / [class] token
An extra token whose output vector stands for the whole input. BERT invented it; ViT reused it.
Vision (ViT)
- Vision Transformer (ViT)
Cut a picture into 16×16 patches, treat them as words, and feed them to an unchanged Transformer encoder.
- Images as patches
A 224×224 picture becomes 196 patches of 768 numbers each: the image's tokens.
- Inductive bias versus data
Built-in assumptions help with little data; with 300 million images, learning from data wins.
- Convolutional networks (ResNet)
The design that ruled computer vision from 2012 to 2020: small filters slid across the picture, with residual shortcuts.
LLM inference
- Prefill and decode
How an LLM answers: read the whole prompt at once, then write one token at a time.
- KV cache
The keys and values of every past token, kept in memory so they are not recomputed. It is the biggest memory bill.
- vLLM: paging and batching
Serving thousands of requests by paging the KV cache like an operating system and batching continuously.
- Speculative decoding
A small model drafts several tokens; the big model checks them in one step. Same words, fewer steps.
Retrieval and RAG
- RAG
Retrieve the right passages first, then let the model answer from them. A full book, from first principles.
- Embeddings
Meaning as geometry: texts become vectors, and similar meanings land close together.
- Vector search and HNSW
Finding the nearest vectors among millions fast, with a layered graph of neighbours.
- BM25 and hybrid search
Keyword scores and vector scores, fused, beat either alone.
- Rerankers
A second, slower model (often BERT-like) re-orders the top results more carefully.
- Evaluating RAG
Every retrieval and generation metric, and how to measure a pipeline honestly.
Agents
- LangGraph and agentic RAG
RAG as a graph of steps that can loop, branch and call tools.
- Multi-agent systems and guardrails
Several agents working together, and the checks that keep them in bounds.
- Jev: a model that decides
A video on an AI model that does not talk, it decides.
GPUs and hardware
- Why GPUs exist
Thousands of simple cores doing the same arithmetic at once: exactly what matrix multiplication needs.
- The CUDA programming model
Threads, blocks and grids: how work is handed to a GPU.
- CPU parallelism
What a CPU can and cannot do in parallel, and why that is not enough.