GPU Optimization
Find the limiter, then fix the limiter. A map of the techniques and when each one applies.
Almost every GPU optimization helps in exactly one of two situations: when the arithmetic units are waiting on memory, or when memory is keeping up and the arithmetic is the constraint. Applying a technique to the wrong one wastes the afternoon, which is why the first page below is about deciding which you are in rather than about making anything faster.
After that, work down the table. The foundational rows apply to almost any kernel and are where the large, cheap wins live. The advanced rows are situational and are where the remaining wins live once the basics are done.
| Technique | Helps when | Page |
|---|---|---|
| Classify the kernel first | Always, before anything else | The Optimization Mindset |
| Measure it properly | Always | GPU Profiling Tools |
| Arithmetic intensity and the ridge point | You want the decision made numerically | The Roofline Model |
| Coalesce global access | Memory-bound, scattered access pattern | Memory Coalescing |
| Raise or lower occupancy | Latency-bound, nothing ready to issue | Occupancy and Register Pressure |
| Avoid divergent branches | Lanes in a group disagree in a hot loop | Warps vs Wavefronts |
| Fuse kernels | Memory-bound chains of small operations | Kernel Fusion |
| Narrow the precision | Either regime, for different reasons | Mixed Precision |
| Tile for reuse | Compute-heavy work that keeps re-reading operands | Tiling and Blocking |
| Pad the scratchpad stride | Heavy shared memory traffic in the inner loop | Shared Memory and Bank Conflicts |
| Overlap copies with compute | Meaningful host-device transfer alongside work | Streams |
| Search the parameter space | Hand-tuning has stopped improving things | Autotuning and Kernel Generation |
Divergence is an execution-model property rather than a technique, so it is explained where the execution model is, on Warps vs Wavefronts, including what a divergent branch actually costs and why it is a data layout problem rather than a control flow one.
Latency hiding across operations is the same mechanism as overlapping transfers with computation, so it lives on Streams, together with the pinned memory requirement that silently defeats it. Latency hiding within a kernel is occupancy, which has its own page.
Profile the whole program before any kernel, because the kernel you assume dominates frequently does not. Confirm the device is actually busy; a low utilisation number means the problem is upstream and no kernel change will help.
Then take the dominant kernel, establish whether it is limited by bandwidth or arithmetic, and pick from the matching half of the table. Change one thing at a time and measure outside the profiler, since a kernel profiler replays kernels and its wall clock is not your wall clock.
Stop when the kernel is close to the relevant ceiling. A kernel at a high fraction of peak bandwidth cannot be improved without changing the algorithm, and recognising that is worth as much as any technique here.