the short version

There are two reasons a GPU kernel is slow, and they have opposite cures. Either the arithmetic units are waiting for data, or the data is waiting for arithmetic units. Optimizations that help one do nothing for the other, and often make it slightly worse.

So the first question is never "how do I make this faster". It is "which of the two am I". Answer that before you change a line.

the two regimes

Memory-boundCompute-bound
The limit isBytes per second off HBMArithmetic throughput
SymptomHigh bandwidth utilisation, low arithmetic utilisationThe reverse
Typical kernelsElementwise ops, normalization, softmax, token-by-token attentionLarge GEMMs, convolutions
What helpsFusion, better access patterns, lower precision, more reuseBetter tiling, matrix units, instruction scheduling
What does nothingFaster math, matrix unitsFusing, reducing traffic

Most kernels people write are in the left column, and most optimization advice on the internet is written for the right one. That mismatch is the single most common reason an afternoon of work produces nothing.

the gut check

Before reaching for a profiler, you can usually guess correctly in a few seconds by asking how much arithmetic happens per byte loaded.

If each element is read, one or two operations happen, and it is written back, you are memory-bound. Vector addition, activation functions, normalization and elementwise scaling are all firmly in this category, and no amount of cleverness changes that.

If each byte loaded participates in many operations, you are potentially compute-bound. A matrix multiply reads each input element and uses it across an entire row or column of output, so it has enormous reuse available. Whether you actually achieve that reuse is the whole subject of tiling.

The ratio has a name, arithmetic intensity, and the roofline model turns the gut check into an actual number with an actual threshold.

then measure anyway

The gut check tells you where to look. It does not tell you what is happening, and GPU intuition is wrong often enough to be worth distrusting on principle.

Two numbers settle it: achieved memory bandwidth as a fraction of the device peak, and achieved arithmetic throughput as a fraction of its peak. Whichever is closer to its ceiling is your limiter. A kernel at 85% of bandwidth and 4% of peak FLOPs is not a kernel with a compute problem, no matter how much arithmetic appears in the source.

If neither is close to its ceiling, you have a third problem, usually latency: not enough parallelism in flight to keep the machine busy, or too many synchronization points draining it.

the discipline

Measure before and after every change, separately. Bundling five optimizations and observing that the total got faster teaches you nothing about which of the five mattered, and one of them is often making things worse.

Measure correctly. A GPU launch is asynchronous, so timing it with a host clock around the launch measures the enqueue and not the kernel. Warm up first, since the first call pays compilation and allocation costs, and average several runs because clocks vary with temperature.

Fix the biggest thing first. If a kernel is 90% of your runtime, a 2x improvement there beats eliminating everything else entirely. Profile the whole program before optimizing any part of it.

Know when to stop. A kernel at a high fraction of the relevant ceiling cannot be improved further without changing the algorithm. Recognising that is worth as much as any optimization.

see it on your own machine

ncu --set full ./my_app | https://docs.nvidia.com/nsight-compute/ | per-kernel achieved bandwidth and arithmetic throughput, which together answer the question this page is about |'om_ncu'
nsys profile --stats=true ./my_app | https://docs.nvidia.com/nsight-systems/ | whole-program timeline first, to find out which kernel is worth optimizing at all |'om_nsys'
rocprofv3 --kernel-trace --stats -- ./my_app | https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/ | the AMD starting point for the same two questions |'om_rocprof'

related topics

The Roofline Model in Practice — turning the gut check into a number and a threshold.
Profiling Tools — which profiler answers which question.
CPU vs GPU: Why GPUs Exist — why bandwidth is so often the binding constraint.

reference

NVIDIA Nsight Compute
AMD rocprofiler
CUDA C++ Best Practices Guide