Occupancy and Register Pressure
More resident threads means more latency hidden. It also means fewer registers each, and there is a point where that stops paying.
Occupancy is the ratio of thread groups actually resident on a compute unit to the maximum that unit could hold. It matters because residency is how the hardware hides memory latency: when one group stalls on a load, the scheduler issues from another. No spare groups means nothing to switch to, and the unit idles.
It is also the most over-optimized number in GPU programming. Higher is not automatically better, and chasing 100% has made plenty of kernels slower.
Each compute unit has a fixed supply of three things, and whichever runs out first sets your occupancy.
| Resource | How a block consumes it | What it caps |
|---|---|---|
| Registers | Registers per thread × threads per block | Usually the binding limit in a serious kernel |
| Shared memory / LDS | Whatever the block statically or dynamically allocates | The binding limit in tiled kernels |
| Block and thread slots | A hard architectural maximum per unit | Rarely binding, but it exists |
This is why occupancy is a compile-time consequence of choices inside your kernel, not a runtime setting. Adding a few local variables can push registers per thread over a threshold, drop resident blocks from six to four, and cost you performance in a way that looks unrelated to the change you made.
The compiler assigns each thread's locals to registers. When it cannot fit within the per-thread budget, it spills the excess to device memory, which is backed by cache but is still hundreds of times slower than a register.
Spilling is the failure mode worth watching, because it is invisible in the source and catastrophic in effect. A kernel that spills in its inner loop can be several times slower than the same kernel restructured to use fewer live values, with identical arithmetic.
You can also cap registers explicitly, either with a compiler flag or with launch bounds in the source, which tells the compiler how many threads per block you intend so it can budget accordingly. Capping too aggressively just moves the cost into spills, so this is a measure-and-compare exercise rather than a setting to copy.
Latency hiding has diminishing returns. Once there are enough resident groups that the scheduler almost always has something ready, adding more buys nothing, and the registers you gave up to get there were doing useful work.
The opposite strategy is often better: give each thread more registers so it can hold more values and do more arithmetic per load, accept lower occupancy, and hide latency through instruction-level parallelism within each thread instead. Well-tuned matrix multiply kernels typically run at modest occupancy for exactly this reason, and beat higher-occupancy versions of themselves.
The practical rule: treat low occupancy as a hypothesis to investigate, not a bug to fix. If a kernel is at 25% occupancy and also at 85% of peak bandwidth, occupancy is not your problem and raising it will not help.
Occupancy is worth raising when the profiler shows the unit stalling on memory with no eligible group to issue from, and your arithmetic and bandwidth utilisation are both low. That combination means the machine is waiting with nothing to do, which is exactly what residency fixes.
The levers, in rough order of how often they work: reduce registers per thread by shortening the lifetime of locals, reduce shared memory per block by tiling smaller, change the block size so blocks pack into the unit more efficiently, or split one large kernel into two smaller ones with lighter resource demands.