GPU Programming
From zero to kernel engineer
A complete course that teaches you to predict how fast a kernel can go, measure how fast it does go, and close the gap. From CPU parallelism to Hopper, Blackwell, FlashAttention and multi-GPU inference.
A GPU kernel is fast when it keeps the machine's bottleneck resource busy. There are only a handful of candidates (memory bandwidth, memory latency, compute throughput, instruction issue, synchronisation, launch overhead and interconnect), and this book teaches you to identify which one you are fighting and what to do about it.
Every chapter ends with the kind of interview questions asked for kernel and performance roles, with answers.
Running the code
Chapter 1 runs on any laptop CPU. From Chapter 2 on you need an NVIDIA GPU; a free Google Colab T4 is enough. The code lives in code/gpu in this site's repository.
Chapters
- Chapter 1 · CPU Parallelism: What GPUs Are Reacting Against
The reference machine for this chapter is an Apple M5 Pro laptop (MacBook Pro, Mac17,8, macOS 26.5.1, 64 GB). Apple's published facts: an 18-core CPU made of 6 "super" cores and 12 "performance" cores, up to 307 GB/s…
- Chapter 2 · Why GPUs Exist
GPUs were built to color pixels. Each pixel is independent of the others: thousands of identical tiny programs. In the mid-2000s people noticed that a matrix multiply looks the same: thousands of independent dot…
- Chapter 3 · The Hardware
You can ignore GPCs and TPCs: they're a graphics-era grouping (and the number of SMs per TPC is not always 2). The SM is the unit you program against.
- Chapter 4 · The Programming Model
You write a function, called a kernel, that describes the work of one thread. You then launch many thousands of copies of it.
- Chapter 5 · The Toolchain
nvcc is not a compiler. It's a driver that splits your .cu file in two and hands each half to a different compiler.