The Roofline Model in Practice
One number decides which half of the optimization advice applies to you.
The roofline model turns "is this memory-bound or compute-bound" from a judgement call into arithmetic. You compute one ratio for your kernel, compare it against one ratio for your hardware, and the comparison tells you which optimizations can possibly help.
It is the most useful back-of-envelope model in GPU work precisely because it rules things out. Knowing an optimization cannot help is worth more than a vague sense that it might.
Arithmetic intensity is floating point operations performed, divided by bytes moved to and from device memory. It is a property of the algorithm and its implementation, not of the hardware.
| Kernel | Roughly | Intensity |
|---|---|---|
| Vector add, FP32 | 2 reads and 1 write per 1 add | About 0.08 FLOP/byte |
| Elementwise activation | 1 read and 1 write per few ops | Similarly tiny |
| Matmul, naive | Re-reads operands for every output | Low, despite the arithmetic |
| Matmul, well tiled | Each element loaded once and reused across a tile | High, and rising with tile size |
| Attention, long sequence | Dominated by moving the score matrix | Low until it is fused |
The two matmul rows matter most. Identical arithmetic, wildly different intensity, purely because of how much reuse the implementation achieves before data leaves the chip. Intensity is something you change, not just something you measure.
The hardware supplies the other number: peak arithmetic throughput divided by peak memory bandwidth. That ratio is the ridge point, the intensity at which a kernel could in principle saturate both at once.
Below the ridge, you cannot feed the arithmetic units fast enough and your ceiling is bandwidth times your intensity. Above it, memory keeps up and your ceiling is peak arithmetic. Plotting achievable performance against intensity gives a line that rises then flattens, which is where the name comes from.
Modern accelerators have high ridge points, in the hundreds of FLOPs per byte for matrix math, because arithmetic throughput has grown faster than bandwidth for years. That leaves a lot of work on the sloped part of the roof: elementwise ops, normalization, reductions, and LLM decode at small batch sizes. It does not leave everything there. A large, well-tiled matrix multiply clears the ridge comfortably, which is why training and prompt prefill are usually compute-bound. The next section shows how to tell from the spec sheet which side a given kernel lands on.
You do not need a profiler to draw a roofline. Two rows of the spec sheet are enough, and working it through by hand once makes the chart a profiler draws much easier to trust. The running example is an AMD MI300X.
1. Pick the right two numbers. Peak memory bandwidth is one number per card. Peak arithmetic is not: there is a separate figure for each precision, and for matrix units versus ordinary vector ALUs. Use the one your kernel actually runs on, and the dense figure rather than the sparse one, which assumes a structured-sparsity pattern your data almost certainly does not have.
| MI300X roof | Peak arithmetic | Bandwidth | Ridge point |
|---|---|---|---|
| BF16 on the matrix cores, dense | 1307 TFLOPS | 5.3 TB/s | 1307 ÷ 5.3 ≈ 247 FLOP/byte |
| FP32 on the vector ALUs | 163 TFLOPS | 5.3 TB/s | 163 ÷ 5.3 ≈ 31 FLOP/byte |
2. Write the roof down. Attainable performance at intensity I is min(peak FLOPS, bandwidth × I). Evaluating it at a few points is all the chart is:
| Intensity (FLOP/byte) | 5.3 TB/s × I | Attainable, BF16 roof | Limited by |
|---|---|---|---|
| 1 | 5.3 TFLOPS | 5.3 TFLOPS | Memory |
| 10 | 53 TFLOPS | 53 TFLOPS | Memory |
| 100 | 530 TFLOPS | 530 TFLOPS | Memory |
| 247 | 1307 TFLOPS | 1307 TFLOPS | Both: the ridge |
| 1000 | 5300 TFLOPS | 1307 TFLOPS | Compute |
3. One roof per precision. Every roof on a card shares the same sloped part, because they all draw on the same memory. Only the flat part moves, and the ridge moves with it. On an MI300X, FP32 vector code turns compute-bound at an intensity of about 31, eight times sooner than BF16 matrix code. Narrowing the precision also raises intensity, since each value costs fewer bytes, so it moves the kernel and the roof at the same time.
4. Place the kernel. Count FLOPs and bytes for the work, assuming each operand crosses the memory bus once, and read off where it lands. The four kernels on the chart, all in BF16:
| Kernel | FLOPs | Bytes | Intensity | Ceiling on an MI300X |
|---|---|---|---|---|
| Vector add, n elements | n | 6n: two 2-byte reads, one 2-byte write | ≈ 0.17 | ≈ 0.9 TFLOPS, memory-bound |
| LLM decode, batch 1, N×N weights | 2N²: a multiply and an add per weight | ≈ 2N²: every weight read once | ≈ 1 | ≈ 5.3 TFLOPS, under 1% of peak |
| The same layer, batch of B | 2BN² | ≈ 2N²: weights still read once | ≈ B | B = 64 gives ≈ 340 TFLOPS; crosses the ridge near B ≈ 250 |
| Square matmul, N×N×N | 2N³ | 6N²: A and B read, C written | N/3 | Compute-bound past N ≈ 740; N = 4096 sits at ≈ 1365 |
The decode rows are the clearest picture of why batching matters in LLM serving: the weights cost the same bytes whether one token or a hundred uses them, so each extra token in the batch is close to free until the batch reaches the ridge. The matmul row is the ideal case. It assumes perfect reuse, which only a well-tiled kernel gets near; a naive one re-reads its operands and lands far to the left.
5. Redraw it with measured ceilings. The spec sheet gives you the ceiling in theory. For a chart you can plan against, measure both numbers on your own card: a large streaming copy (BabelStream is the usual benchmark) for bandwidth, and a large GEMM from rocBLAS or hipBLASLt (cuBLAS on Nvidia) for arithmetic. Both usually come in noticeably under the printed figures, which lowers both roofs and shifts the ridge point.
Place your kernel at its measured intensity, at its measured performance. Three cases follow.
On the sloped roof. You are bandwidth-limited and achieving close to what bandwidth allows. Adding compute, matrix units or faster math does nothing. To go faster you must move fewer bytes: fuse, improve reuse, or narrow the precision.
On the flat roof. You are compute-limited and near peak. Reducing memory traffic buys nothing. Improve the arithmetic: reach the matrix units, narrow the precision, schedule better.
Well below both. The common and least satisfying case. Neither resource is saturated, so something else is the problem: not enough parallelism in flight, too many synchronization points, divergence, or a latency chain you have not spotted. The roofline tells you to stop looking at bandwidth and arithmetic entirely.
It is good for deciding what not to try, for sanity-checking a claimed speedup, and for setting a realistic target before you start. If a kernel sits at 90% of its bandwidth roof, the remaining 10% is the whole prize, and knowing that stops you spending a week on it.
It is a simplification. It assumes traffic goes to device memory, so caches make real kernels beat the naive prediction. It ignores latency, occupancy and instruction mix. It treats a kernel as one homogeneous thing when phases may behave differently. Treat it as a bound and a direction, not a prediction.
Nsight Compute draws the chart for you and places the kernel on it, which is by far the easiest way in.