The Optimization Mindset: Measure First
Decide whether the kernel is waiting on memory or on arithmetic. Everything else follows from that one answer.
There are two reasons a GPU kernel is slow, and they have opposite cures. Either the arithmetic units are waiting for data, or the data is waiting for arithmetic units. Optimizations that help one do nothing for the other, and often make it slightly worse.
So the first question is never "how do I make this faster". It is "which of the two am I". Answer that before you change a line.
| Memory-bound | Compute-bound | |
|---|---|---|
| The limit is | Bytes per second off HBM | Arithmetic throughput |
| Symptom | High bandwidth utilisation, low arithmetic utilisation | The reverse |
| Typical kernels | Elementwise ops, normalization, softmax, token-by-token attention | Large GEMMs, convolutions |
| What helps | Fusion, better access patterns, lower precision, more reuse | Better tiling, matrix units, instruction scheduling |
| What does nothing | Faster math, matrix units | Fusing, reducing traffic |
Most kernels people write are in the left column, and most optimization advice on the internet is written for the right one. That mismatch is the single most common reason an afternoon of work produces nothing.
Before reaching for a profiler, you can usually guess correctly in a few seconds by asking how much arithmetic happens per byte loaded.
If each element is read, one or two operations happen, and it is written back, you are memory-bound. Vector addition, activation functions, normalization and elementwise scaling are all firmly in this category, and no amount of cleverness changes that.
If each byte loaded participates in many operations, you are potentially compute-bound. A matrix multiply reads each input element and uses it across an entire row or column of output, so it has enormous reuse available. Whether you actually achieve that reuse is the whole subject of tiling.
The ratio has a name, arithmetic intensity, and the roofline model turns the gut check into an actual number with an actual threshold.
The gut check tells you where to look. It does not tell you what is happening, and GPU intuition is wrong often enough to be worth distrusting on principle.
Two numbers settle it: achieved memory bandwidth as a fraction of the device peak, and achieved arithmetic throughput as a fraction of its peak. Whichever is closer to its ceiling is your limiter. A kernel at 85% of bandwidth and 4% of peak FLOPs is not a kernel with a compute problem, no matter how much arithmetic appears in the source.
If neither is close to its ceiling, you have a third problem, usually latency: not enough parallelism in flight to keep the machine busy, or too many synchronization points draining it.
Measure before and after every change, separately. Bundling five optimizations and observing that the total got faster teaches you nothing about which of the five mattered, and one of them is often making things worse.
Measure correctly. A GPU launch is asynchronous, so timing it with a host clock around the launch measures the enqueue and not the kernel. Warm up first, since the first call pays compilation and allocation costs, and average several runs because clocks vary with temperature.
Fix the biggest thing first. If a kernel is 90% of your runtime, a 2x improvement there beats eliminating everything else entirely. Profile the whole program before optimizing any part of it.
Know when to stop. A kernel at a high fraction of the relevant ceiling cannot be improved further without changing the algorithm. Recognising that is worth as much as any optimization.