Writings
- SGLang Explained: How It Reuses Work, and How It Differs from vLLM
Follow a handbook assistant from its first request to a shared KV cache. Learn RadixAttention step by step, then compare SGLang with vLLM.
- Attention, Part 1: Self-Attention and Multi-Head Attention
The one idea every modern AI language model is built on, explained from zero: why attention was invented, how a word decides which other words to look at, why the scores are divided by √d, how several heads work together, and what real attention heads inside Qwen2.5 actually learn. Every claim checked in code.
- Attention, Part 2: MQA, GQA and MLA
Every token leaves its keys and values behind in memory (the KV cache), and plain multi-head attention leaves a lot. How multi-query, grouped-query and multi-head latent attention shrink it, explained from zero, built from scratch, checked against a real model's layer, and measured: DeepSeek-V3 stores 68.6 KiB per token where plain attention would need 4,880.
- Attention, Part 3: Sliding-Window and Sparse Attention
Why looking at every earlier token gets too expensive, and the three ways today's models avoid it: sliding windows (Mistral), mixing local and global layers (Gemma 2, Gemma 3, gpt-oss), and letting the model pick its own tokens (DeepSeek Sparse Attention). Explained from zero, with equations, quotes from the papers, and real experiments: a lightning indexer trained on Qwen2.5-0.5B keeps quality close to full attention while each token reads only 64 of up to 2,048 earlier tokens.
- Attention, Part 4: Linear Attention, Gated DeltaNet and Hybrid Models
The newest idea in attention: replace the ever-growing KV cache with a fixed-size memory. Linear attention, the delta rule, forget gates, Gated DeltaNet, Kimi Delta Attention, gated attention, and the hybrid models built from them (Qwen3-Next, Kimi Linear, Nemotron-H). Explained from zero with equations, paper screenshots and real runs, including three small models trained side by side.
- How an LLM Writes: Prefill and Decode
Why the first word takes a moment and the rest stream steadily. The two phases of LLM inference, built from the ground up and measured on a real model: a 512-token prompt is read in 27 ms, but writing 512 tokens takes 5.2 seconds.
- The KV Cache: An LLM's Short-Term Memory, and Its Bill
Why every LLM keeps notes on each token it has read, how big those notes get (16 GiB for one 128K-token Llama 3.1 8B conversation), how models shrink them, and why they, not compute, usually decide how many people one GPU can serve.
- How vLLM Serves Thousands: Paging, Batching and Caching
PagedAttention, continuous batching, chunked prefill and prefix caching, each built from the problem it solves and measured or simulated with real costs: from 48 to 415 requests on one GPU, pauses cut from 463 ms to 40 ms, and a 7x faster first token.
- Speculative Decoding: Several Tokens per Step, Not One Word Changed
How a small model can guess ahead and a big model can check all the guesses at once, why the output stays exactly the same, and what it really buys on real hardware: speculative decoding built from scratch and measured, from 1.16x on prose to 1.89x on code and 3.49x when the answer copies the prompt.
- HNSW, From the Ground Up
How Hierarchical Navigable Small World graphs find nearest neighbours in microseconds: the two ideas they combine, how the graph is built, what M, efConstruction and efSearch really cost, and measured numbers from a 200,000-vector index.
- How this site renders Markdown
A tour of every element the reader supports (headings, code, callouts, tables, math and diagrams), so writing a new chapter is just dropping a .md file into content/.