the short version

Occupancy is the ratio of thread groups actually resident on a compute unit to the maximum that unit could hold. It matters because residency is how the hardware hides memory latency: when one group stalls on a load, the scheduler issues from another. No spare groups means nothing to switch to, and the unit idles.

It is also the most over-optimized number in GPU programming. Higher is not automatically better, and chasing 100% has made plenty of kernels slower.

what caps it

Each compute unit has a fixed supply of three things, and whichever runs out first sets your occupancy.

ResourceHow a block consumes itWhat it caps
RegistersRegisters per thread × threads per blockUsually the binding limit in a serious kernel
Shared memory / LDSWhatever the block statically or dynamically allocatesThe binding limit in tiled kernels
Block and thread slotsA hard architectural maximum per unitRarely binding, but it exists

This is why occupancy is a compile-time consequence of choices inside your kernel, not a runtime setting. Adding a few local variables can push registers per thread over a threshold, drop resident blocks from six to four, and cost you performance in a way that looks unrelated to the change you made.

register pressure and spilling

The compiler assigns each thread's locals to registers. When it cannot fit within the per-thread budget, it spills the excess to device memory, which is backed by cache but is still hundreds of times slower than a register.

Spilling is the failure mode worth watching, because it is invisible in the source and catastrophic in effect. A kernel that spills in its inner loop can be several times slower than the same kernel restructured to use fewer live values, with identical arithmetic.

nvcc --ptxas-options=-v -O3 kernel.cu -o kernel | https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/ | prints registers per thread, shared memory per block, and spill stores and loads in bytes; nonzero spill counts in a hot kernel are worth chasing |'oc_ptxas'

You can also cap registers explicitly, either with a compiler flag or with launch bounds in the source, which tells the compiler how many threads per block you intend so it can budget accordingly. Capping too aggressively just moves the cost into spills, so this is a measure-and-compare exercise rather than a setting to copy.

why maximum occupancy is not the goal

Latency hiding has diminishing returns. Once there are enough resident groups that the scheduler almost always has something ready, adding more buys nothing, and the registers you gave up to get there were doing useful work.

The opposite strategy is often better: give each thread more registers so it can hold more values and do more arithmetic per load, accept lower occupancy, and hide latency through instruction-level parallelism within each thread instead. Well-tuned matrix multiply kernels typically run at modest occupancy for exactly this reason, and beat higher-occupancy versions of themselves.

The practical rule: treat low occupancy as a hypothesis to investigate, not a bug to fix. If a kernel is at 25% occupancy and also at 85% of peak bandwidth, occupancy is not your problem and raising it will not help.

when it genuinely is the problem

Occupancy is worth raising when the profiler shows the unit stalling on memory with no eligible group to issue from, and your arithmetic and bandwidth utilisation are both low. That combination means the machine is waiting with nothing to do, which is exactly what residency fixes.

The levers, in rough order of how often they work: reduce registers per thread by shortening the lifetime of locals, reduce shared memory per block by tiling smaller, change the block size so blocks pack into the unit more efficiently, or split one large kernel into two smaller ones with lighter resource demands.

related topics

SM vs CU — the resource budget occupancy is computed against.
Warps vs Wavefronts — the groups being counted.
The Optimization Mindset — deciding whether occupancy is even your problem.

reference

CUDA C++ Best Practices Guide
CUDA Occupancy Calculator
AMD HIP documentation