The GPU Memory Hierarchy
Registers to HBM in five steps, each one bigger and roughly ten times slower.
GPU memory is a ladder. At the top are registers, which are effectively free to read and tiny. At the bottom is device memory, which holds everything and takes hundreds of cycles to reach. In between are a programmer-managed scratchpad and two levels of automatic cache. Registers, the scratchpad and L1 belong to a single SM (Nvidia) or CU (AMD); L2 and device memory are shared more widely.
Every GPU optimization worth the name is the same move: get data up this ladder once, then do as much work as possible before it falls back down.
| Level | Managed by | Visible to | Rough size |
|---|---|---|---|
| Registers | The compiler | One thread | Hundreds of KB per SM/CU, in aggregate |
| Shared memory / LDS | You, explicitly | All threads in one block | Tens to a couple hundred KB per block |
| L1 / texture cache | Hardware | One SM/CU | Tens to a couple hundred KB per SM/CU |
| L2 cache | Hardware | The whole die | Tens of MB |
| HBM / GDDR | You, via allocations | The whole device, and the host | Gigabytes to hundreds of GB |
One structural difference worth knowing: AMD's recent datacenter parts insert a large last-level Infinity Cache between L2 and memory, so the ladder there has an extra rung that absorbs a good deal of traffic before it reaches HBM.
The register file is, surprisingly, the largest on-chip memory on the machine, larger than the shared memory and typically larger than L1. It has to be, because it holds the private state of every resident thread simultaneously, and an SM/CU may keep thousands of threads resident.
That is also where the tension lives. Registers per SM/CU are fixed, so registers per thread multiplied by resident threads cannot exceed the supply. Using more registers per thread makes each thread faster and reduces how many threads fit, which reduces the machine's ability to hide memory latency. When the compiler cannot fit, it spills to device memory, and a spilled kernel can be dramatically slower than the register pressure suggests. This trade is the whole subject of occupancy tuning.
Shared memory on Nvidia, LDS on AMD, is a scratchpad carved out per block and visible to every thread in that block. It is the only level you place data in deliberately, and it is the reason a block is assigned to exactly one SM/CU: the scratchpad is physically part of that SM/CU.
The canonical use is tiling. Rather than have every thread read its operands from device memory, one block cooperatively loads a tile into the scratchpad, synchronizes, and then every thread reads its operands from there many times over. A well-tiled matrix multiply reads each input element from device memory roughly once instead of hundreds of times, and that single change is usually worth more than every other optimization combined.
The scratchpad is banked, which introduces its own failure mode when many lanes hit the same bank at once. That is a page of its own later in the optimization section.
L1 sits inside an SM/CU and on Nvidia shares physical storage with shared memory, so the split between them is configurable. L2 is shared by the entire die and is the last stop before device memory, which makes it the coherence point between SMs/CUs and the thing that quietly saves you when several blocks happen to read the same data.
Because they are automatic, the useful question is not how to use them but whether your access pattern lets them work. A kernel whose threads read neighbouring addresses gets full value from every cache line fetched. One whose threads read scattered addresses pulls in full lines and uses a fraction of each, and the hierarchy cannot save it.
Put the two ends together. Device memory delivers enormous bandwidth but takes hundreds of cycles to answer, while the arithmetic units can consume operands every cycle. The gap between those two numbers is the central problem of GPU programming, and the hierarchy exists to close it.
That gives a single question to ask of any kernel: how much arithmetic do you do per byte you pull off the device? If the answer is "very little", no amount of cleverness inside the kernel will help, because you are waiting on memory and the arithmetic units are idle. If the answer is "a lot", the tiling and reuse this page describes is exactly where your speedup lives. That ratio has a name, arithmetic intensity, and a model built around it.
The limits per SM/CU are queryable, and the profilers report what fraction of your traffic each level actually served.