Topic map

Attention and Transformers

Pre-training (BERT)

  • BERT

    A Transformer encoder that reads text in both directions, pre-trained once and fine-tuned for any task.

  • Masked language modelling

    Hide 15% of the words and train the model to guess them from both sides.

  • Pre-train, then fine-tune

    Train once on a huge dataset, then adapt a copy to each task with one small new layer.

  • The [CLS] / [class] token

    An extra token whose output vector stands for the whole input. BERT invented it; ViT reused it.

Vision (ViT)

  • Vision Transformer (ViT)

    Cut a picture into 16×16 patches, treat them as words, and feed them to an unchanged Transformer encoder.

  • Images as patches

    A 224×224 picture becomes 196 patches of 768 numbers each: the image's tokens.

  • Inductive bias versus data

    Built-in assumptions help with little data; with 300 million images, learning from data wins.

  • Convolutional networks (ResNet)

    The design that ruled computer vision from 2012 to 2020: small filters slid across the picture, with residual shortcuts.

LLM inference

  • Prefill and decode

    How an LLM answers: read the whole prompt at once, then write one token at a time.

  • KV cache

    The keys and values of every past token, kept in memory so they are not recomputed. It is the biggest memory bill.

  • vLLM: paging and batching

    Serving thousands of requests by paging the KV cache like an operating system and batching continuously.

  • Speculative decoding

    A small model drafts several tokens; the big model checks them in one step. Same words, fewer steps.

Retrieval and RAG

  • RAG

    Retrieve the right passages first, then let the model answer from them. A full book, from first principles.

  • Embeddings

    Meaning as geometry: texts become vectors, and similar meanings land close together.

  • Vector search and HNSW

    Finding the nearest vectors among millions fast, with a layered graph of neighbours.

  • BM25 and hybrid search

    Keyword scores and vector scores, fused, beat either alone.

  • Rerankers

    A second, slower model (often BERT-like) re-orders the top results more carefully.

  • Evaluating RAG

    Every retrieval and generation metric, and how to measure a pipeline honestly.

Agents

GPUs and hardware

  • Why GPUs exist

    Thousands of simple cores doing the same arithmetic at once: exactly what matrix multiplication needs.

  • The CUDA programming model

    Threads, blocks and grids: how work is handed to a GPU.

  • CPU parallelism

    What a CPU can and cannot do in parallel, and why that is not enough.