<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel>
<title>Ishwar Jangid: writings and research papers</title><link>https://ishwarj.com/</link><description>Ishwar Jangid writes long-form, first-principles guides and books on LLM inference, retrieval-augmented generation, GPU programming and ML systems, with code you can run and numbers you can measure.</description><language>en</language>
<atom:link href="https://ishwarj.com/rss.xml" rel="self" type="application/rss+xml" />
<item><title>SGLang Explained: How It Reuses Work, and How It Differs from vLLM</title><link>https://ishwarj.com/writings/llm-inference-5-sglang-vs-vllm/</link><guid>https://ishwarj.com/writings/llm-inference-5-sglang-vs-vllm/</guid><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><description>Follow a handbook assistant from its first request to a shared KV cache. Learn RadixAttention step by step, then compare SGLang with vLLM.</description></item>
<item><title>ViT, Part 1: The Big Idea: An Image Is Worth 16×16 Words</title><link>https://ishwarj.com/papers/vit/part-1-the-big-idea/</link><guid>https://ishwarj.com/papers/vit/part-1-the-big-idea/</guid><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><description>The title, the abstract and the introduction of the Vision Transformer paper, line by line: what it means to cut a picture into 16×16 patches and read them like words, why convolutional networks ruled computer vision, what inductive bias, locality and translation equivariance are, and the claim that enough data beats built-in assumptions. With a real run of ViT-B/16 and ResNet-50 on the same picture and the quadratic cost of attention worked out by hand.</description></item>
<item><title>ViT, Part 2: Inside ViT: Patches, Embeddings and the Encoder</title><link>https://ishwarj.com/papers/vit/part-2-the-model/</link><guid>https://ishwarj.com/papers/vit/part-2-the-model/</guid><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><description>Section 2, Figure 1, Section 3.1 and Appendix A, line by line: the earlier attempts to put attention on images, then every step from a 224×224 picture to a class label, with Equations 1 to 8 recomputed by hand on the released ViT-B/16 and matched against the library at every stage, the 86,567,656 parameters counted exactly, and a four-token attention example small enough to check with a pencil.</description></item>
<item><title>Attention, Part 1: Self-Attention and Multi-Head Attention</title><link>https://ishwarj.com/writings/attention-1-self-attention/</link><guid>https://ishwarj.com/writings/attention-1-self-attention/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>The one idea every modern AI language model is built on, explained from zero: why attention was invented, how a word decides which other words to look at, why the scores are divided by √d, how several heads work together, and what real attention heads inside Qwen2.5 actually learn. Every claim checked in code.</description></item>
<item><title>Attention, Part 2: MQA, GQA and MLA</title><link>https://ishwarj.com/writings/attention-2-mqa-gqa-mla/</link><guid>https://ishwarj.com/writings/attention-2-mqa-gqa-mla/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>Every token leaves its keys and values behind in memory (the KV cache), and plain multi-head attention leaves a lot. How multi-query, grouped-query and multi-head latent attention shrink it, explained from zero, built from scratch, checked against a real model's layer, and measured: DeepSeek-V3 stores 68.6 KiB per token where plain attention would need 4,880.</description></item>
<item><title>Attention, Part 3: Sliding-Window and Sparse Attention</title><link>https://ishwarj.com/writings/attention-3-sliding-window-sparse/</link><guid>https://ishwarj.com/writings/attention-3-sliding-window-sparse/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>Why looking at every earlier token gets too expensive, and the three ways today's models avoid it: sliding windows (Mistral), mixing local and global layers (Gemma 2, Gemma 3, gpt-oss), and letting the model pick its own tokens (DeepSeek Sparse Attention). Explained from zero, with equations, quotes from the papers, and real experiments: a lightning indexer trained on Qwen2.5-0.5B keeps quality close to full attention while each token reads only 64 of up to 2,048 earlier tokens.</description></item>
<item><title>Attention, Part 4: Linear Attention, Gated DeltaNet and Hybrid Models</title><link>https://ishwarj.com/writings/attention-4-linear-hybrid/</link><guid>https://ishwarj.com/writings/attention-4-linear-hybrid/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>The newest idea in attention: replace the ever-growing KV cache with a fixed-size memory. Linear attention, the delta rule, forget gates, Gated DeltaNet, Kimi Delta Attention, gated attention, and the hybrid models built from them (Qwen3-Next, Kimi Linear, Nemotron-H). Explained from zero with equations, paper screenshots and real runs, including three small models trained side by side.</description></item>
<item><title>BERT, Part 1: The Big Idea: Reading in Both Directions</title><link>https://ishwarj.com/papers/bert/part-1-the-big-idea/</link><guid>https://ishwarj.com/papers/bert/part-1-the-big-idea/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>The title, the abstract and the introduction of the BERT paper, line by line: what pre-training is, the two ways to reuse a pre-trained model, why reading only left to right is a real limit, and BERT's fix of filling in blanks. With a real run of GPT-2 and BERT on the same sentences.</description></item>
<item><title>BERT, Part 2: Inside BERT: Architecture and Input</title><link>https://ishwarj.com/papers/bert/part-2-architecture-and-input/</link><guid>https://ishwarj.com/papers/bert/part-2-architecture-and-input/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>Section 2 and the model half of Section 3, line by line: the ideas BERT builds on, the stack of Transformer layers, where the 110 million parameters come from (counted in the real model), WordPiece tokens, [CLS], [SEP], and the three embeddings that are added together for every token.</description></item>
<item><title>BERT, Part 3: Pre-training: Masked LM and Next Sentence Prediction</title><link>https://ishwarj.com/papers/bert/part-3-pre-training/</link><guid>https://ishwarj.com/papers/bert/part-3-pre-training/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>Section 3.1 and Appendices A.1 and A.2 of the BERT paper, line by line: why a normal language model cannot look both ways, the masked language model and its 80/10/10 rule, next sentence prediction, the training data and the full training recipe. Every claim run on the real model or measured in code.</description></item>
<item><title>BERT, Part 4: Fine-tuning and Results: GLUE, SQuAD and SWAG</title><link>https://ishwarj.com/papers/bert/part-4-fine-tuning-and-results/</link><guid>https://ishwarj.com/papers/bert/part-4-fine-tuning-and-results/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>Sections 3.2 and 4 of the BERT paper, line by line: how one pre-trained model becomes a classifier, a question answerer and a multiple-choice solver by adding a tiny output layer, and what the results tables really say. With a real fine-tune of BERT-base on SST-2 and MRPC, and our own span search scored on the full SQuAD dev sets.</description></item>
<item><title>BERT, Part 5: Ablations: What Really Matters</title><link>https://ishwarj.com/papers/bert/part-5-ablations/</link><guid>https://ishwarj.com/papers/bert/part-5-ablations/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>Section 5 and Appendix C of the BERT paper, line by line: which parts of BERT actually cause its gains. Next sentence prediction, bidirectionality, model size, frozen features versus fine-tuning, training length and the 80/10/10 masking recipe, every table row explained, with real parameter counts, a real perplexity and a real re-run of the feature-based NER experiment.</description></item>
<item><title>BERT, Part 6: After BERT: Impact, Limits and Summary</title><link>https://ishwarj.com/papers/bert/part-6-impact-and-summary/</link><guid>https://ishwarj.com/papers/bert/part-6-impact-and-summary/</guid><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><description>The paper's conclusion, what came next (RoBERTa, ALBERT, DistilBERT, ELECTRA, Sentence-BERT and more), BERT's real limits measured on real models, how BERT-style models are used today, and the whole paper summarised on one page.</description></item>
<item><title>How an LLM Writes: Prefill and Decode</title><link>https://ishwarj.com/writings/llm-inference-1-prefill-and-decode/</link><guid>https://ishwarj.com/writings/llm-inference-1-prefill-and-decode/</guid><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><description>Why the first word takes a moment and the rest stream steadily. The two phases of LLM inference, built from the ground up and measured on a real model: a 512-token prompt is read in 27 ms, but writing 512 tokens takes 5.2 seconds.</description></item>
<item><title>The KV Cache: An LLM's Short-Term Memory, and Its Bill</title><link>https://ishwarj.com/writings/llm-inference-2-kv-cache/</link><guid>https://ishwarj.com/writings/llm-inference-2-kv-cache/</guid><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><description>Why every LLM keeps notes on each token it has read, how big those notes get (16 GiB for one 128K-token Llama 3.1 8B conversation), how models shrink them, and why they, not compute, usually decide how many people one GPU can serve.</description></item>
<item><title>How vLLM Serves Thousands: Paging, Batching and Caching</title><link>https://ishwarj.com/writings/llm-inference-3-vllm/</link><guid>https://ishwarj.com/writings/llm-inference-3-vllm/</guid><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><description>PagedAttention, continuous batching, chunked prefill and prefix caching, each built from the problem it solves and measured or simulated with real costs: from 48 to 415 requests on one GPU, pauses cut from 463 ms to 40 ms, and a 7x faster first token.</description></item>
<item><title>Speculative Decoding: Several Tokens per Step, Not One Word Changed</title><link>https://ishwarj.com/writings/llm-inference-4-speculative-decoding/</link><guid>https://ishwarj.com/writings/llm-inference-4-speculative-decoding/</guid><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><description>How a small model can guess ahead and a big model can check all the guesses at once, why the output stays exactly the same, and what it really buys on real hardware: speculative decoding built from scratch and measured, from 1.16x on prose to 1.89x on code and 3.49x when the answer copies the prompt.</description></item>
<item><title>HNSW, From the Ground Up</title><link>https://ishwarj.com/writings/hnsw-from-the-ground-up/</link><guid>https://ishwarj.com/writings/hnsw-from-the-ground-up/</guid><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate><description>How Hierarchical Navigable Small World graphs find nearest neighbours in microseconds: the two ideas they combine, how the graph is built, what M, efConstruction and efSearch really cost, and measured numbers from a 200,000-vector index.</description></item>
<item><title>How this site renders Markdown</title><link>https://ishwarj.com/writings/how-this-site-renders-markdown/</link><guid>https://ishwarj.com/writings/how-this-site-renders-markdown/</guid><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate><description>A tour of every element the reader supports (headings, code, callouts, tables, math and diagrams), so writing a new chapter is just dropping a .md file into content/.</description></item>
</channel></rss>
