the short version

GPU memory is a ladder. At the top are registers, which are effectively free to read and tiny. At the bottom is device memory, which holds everything and takes hundreds of cycles to reach. In between are a programmer-managed scratchpad and two levels of automatic cache. Registers, the scratchpad and L1 belong to a single SM (Nvidia) or CU (AMD); L2 and device memory are shared more widely.

The GPU memory hierarchy Five levels drawn as bars of increasing width: registers, shared memory or LDS, L1 cache, L2 cache, and HBM or GDDR. Each step down is larger in capacity and slower to reach, from about one cycle for a register to several hundred cycles for device memory. Registers ~1 cycle · private to one thread Shared / LDS ~20–30 cycles · shared within a block L1 cache ~30–40 cycles · automatic, per SM/CU L2 cache ~200 cycles · shared by the whole die HBM / GDDR ~400–800 cycles · the whole device
Cycle counts are rough and vary by architecture. The ratios are the durable part: each level down costs roughly an order of magnitude more to reach than the one above it.

Every GPU optimization worth the name is the same move: get data up this ladder once, then do as much work as possible before it falls back down.

the levels, and who can see what

LevelManaged byVisible toRough size
RegistersThe compilerOne threadHundreds of KB per SM/CU, in aggregate
Shared memory / LDSYou, explicitlyAll threads in one blockTens to a couple hundred KB per block
L1 / texture cacheHardwareOne SM/CUTens to a couple hundred KB per SM/CU
L2 cacheHardwareThe whole dieTens of MB
HBM / GDDRYou, via allocationsThe whole device, and the hostGigabytes to hundreds of GB

One structural difference worth knowing: AMD's recent datacenter parts insert a large last-level Infinity Cache between L2 and memory, so the ladder there has an extra rung that absorbs a good deal of traffic before it reaches HBM.

registers are not a small detail

The register file is, surprisingly, the largest on-chip memory on the machine, larger than the shared memory and typically larger than L1. It has to be, because it holds the private state of every resident thread simultaneously, and an SM/CU may keep thousands of threads resident.

That is also where the tension lives. Registers per SM/CU are fixed, so registers per thread multiplied by resident threads cannot exceed the supply. Using more registers per thread makes each thread faster and reduces how many threads fit, which reduces the machine's ability to hide memory latency. When the compiler cannot fit, it spills to device memory, and a spilled kernel can be dramatically slower than the register pressure suggests. This trade is the whole subject of occupancy tuning.

shared memory is the one you control

Shared memory on Nvidia, LDS on AMD, is a scratchpad carved out per block and visible to every thread in that block. It is the only level you place data in deliberately, and it is the reason a block is assigned to exactly one SM/CU: the scratchpad is physically part of that SM/CU.

The canonical use is tiling. Rather than have every thread read its operands from device memory, one block cooperatively loads a tile into the scratchpad, synchronizes, and then every thread reads its operands from there many times over. A well-tiled matrix multiply reads each input element from device memory roughly once instead of hundreds of times, and that single change is usually worth more than every other optimization combined.

The scratchpad is banked, which introduces its own failure mode when many lanes hit the same bank at once. That is a page of its own later in the optimization section.

the caches you do not control

L1 sits inside an SM/CU and on Nvidia shares physical storage with shared memory, so the split between them is configurable. L2 is shared by the entire die and is the last stop before device memory, which makes it the coherence point between SMs/CUs and the thing that quietly saves you when several blocks happen to read the same data.

Because they are automatic, the useful question is not how to use them but whether your access pattern lets them work. A kernel whose threads read neighbouring addresses gets full value from every cache line fetched. One whose threads read scattered addresses pulls in full lines and uses a fraction of each, and the hierarchy cannot save it.

what the ladder is actually for

Put the two ends together. Device memory delivers enormous bandwidth but takes hundreds of cycles to answer, while the arithmetic units can consume operands every cycle. The gap between those two numbers is the central problem of GPU programming, and the hierarchy exists to close it.

That gives a single question to ask of any kernel: how much arithmetic do you do per byte you pull off the device? If the answer is "very little", no amount of cleverness inside the kernel will help, because you are waiting on memory and the arithmetic units are idle. If the answer is "a lot", the tiling and reuse this page describes is exactly where your speedup lives. That ratio has a name, arithmetic intensity, and a model built around it.

see it on your own machine

The limits per SM/CU are queryable, and the profilers report what fraction of your traffic each level actually served.

nvidia-smi --query-gpu=name,memory.total --format=csv | https://developer.nvidia.com/nvidia-system-management-interface | the bottom rung: how much device memory you have to work with |'mh_smi'
rocminfo | https://rocm.docs.amd.com/projects/rocminfo/en/latest/ | reports LDS size per compute unit and the cache line size on AMD hardware |'mh_rocminfo'
ncu --set full ./my_app | https://docs.nvidia.com/nsight-compute/ | per-kernel hit rates for L1 and L2, achieved device memory throughput, and register usage per thread |'mh_ncu'
nvcc --ptxas-options=-v -o app kernel.cu | https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html | prints registers per thread, shared memory per block, and any spilling, at compile time |'mh_ptxas'

related topics

SM vs CU — the unit that owns the register file, the scratchpad and the L1.
Anatomy of a GPU — where HBM physically sits, and why it is fast despite being far away.
GPU Optimization — tiling, coalescing and bank conflicts, as things to measure.

reference

CUDA C++ Programming Guide: device memory accesses
AMD HIP documentation
AMD CDNA architecture
NVIDIA Nsight Compute