how to use this section

Almost every GPU optimization helps in exactly one of two situations: when the arithmetic units are waiting on memory, or when memory is keeping up and the arithmetic is the constraint. Applying a technique to the wrong one wastes the afternoon, which is why the first page below is about deciding which you are in rather than about making anything faster.

After that, work down the table. The foundational rows apply to almost any kernel and are where the large, cheap wins live. The advanced rows are situational and are where the remaining wins live once the basics are done.

the techniques

TechniqueHelps whenPage
Classify the kernel firstAlways, before anything elseThe Optimization Mindset
Measure it properlyAlwaysGPU Profiling Tools
Arithmetic intensity and the ridge pointYou want the decision made numericallyThe Roofline Model
Coalesce global accessMemory-bound, scattered access patternMemory Coalescing
Raise or lower occupancyLatency-bound, nothing ready to issueOccupancy and Register Pressure
Avoid divergent branchesLanes in a group disagree in a hot loopWarps vs Wavefronts
Fuse kernelsMemory-bound chains of small operationsKernel Fusion
Narrow the precisionEither regime, for different reasonsMixed Precision
Tile for reuseCompute-heavy work that keeps re-reading operandsTiling and Blocking
Pad the scratchpad strideHeavy shared memory traffic in the inner loopShared Memory and Bank Conflicts
Overlap copies with computeMeaningful host-device transfer alongside workStreams
Search the parameter spaceHand-tuning has stopped improving thingsAutotuning and Kernel Generation

two that live elsewhere

Divergence is an execution-model property rather than a technique, so it is explained where the execution model is, on Warps vs Wavefronts, including what a divergent branch actually costs and why it is a data layout problem rather than a control flow one.

Latency hiding across operations is the same mechanism as overlapping transfers with computation, so it lives on Streams, together with the pinned memory requirement that silently defeats it. Latency hiding within a kernel is occupancy, which has its own page.

the order that usually works

Profile the whole program before any kernel, because the kernel you assume dominates frequently does not. Confirm the device is actually busy; a low utilisation number means the problem is upstream and no kernel change will help.

Then take the dominant kernel, establish whether it is limited by bandwidth or arithmetic, and pick from the matching half of the table. Change one thing at a time and measure outside the profiler, since a kernel profiler replays kernels and its wall clock is not your wall clock.

Stop when the kernel is close to the relevant ceiling. A kernel at a high fraction of peak bandwidth cannot be improved without changing the algorithm, and recognising that is worth as much as any technique here.

related topics

CPU vs GPU: Why GPUs Exist — why bandwidth is so often the binding constraint.
The GPU Memory Hierarchy — the ladder most of these techniques are moving data up.
Tensor Cores vs Matrix Cores — the units the compute-bound half is trying to reach.
Assembly & Low-Level GPU Systems — what is left once these have been exhausted.

reference

CUDA C++ Best Practices Guide
NVIDIA Nsight Compute
AMD rocprofiler
NVIDIA CUDA C++ Programming Guide