the short version

The roofline model turns "is this memory-bound or compute-bound" from a judgement call into arithmetic. You compute one ratio for your kernel, compare it against one ratio for your hardware, and the comparison tells you which optimizations can possibly help.

It is the most useful back-of-envelope model in GPU work precisely because it rules things out. Knowing an optimization cannot help is worth more than a vague sense that it might.

arithmetic intensity

Arithmetic intensity is floating point operations performed, divided by bytes moved to and from device memory. It is a property of the algorithm and its implementation, not of the hardware.

KernelRoughlyIntensity
Vector add, FP322 reads and 1 write per 1 addAbout 0.08 FLOP/byte
Elementwise activation1 read and 1 write per few opsSimilarly tiny
Matmul, naiveRe-reads operands for every outputLow, despite the arithmetic
Matmul, well tiledEach element loaded once and reused across a tileHigh, and rising with tile size
Attention, long sequenceDominated by moving the score matrixLow until it is fused

The two matmul rows matter most. Identical arithmetic, wildly different intensity, purely because of how much reuse the implementation achieves before data leaves the chip. Intensity is something you change, not just something you measure.

the ridge point

The hardware supplies the other number: peak arithmetic throughput divided by peak memory bandwidth. That ratio is the ridge point, the intensity at which a kernel could in principle saturate both at once.

Below the ridge, you cannot feed the arithmetic units fast enough and your ceiling is bandwidth times your intensity. Above it, memory keeps up and your ceiling is peak arithmetic. Plotting achievable performance against intensity gives a line that rises then flattens, which is where the name comes from.

Modern accelerators have high ridge points, in the hundreds of FLOPs per byte for matrix math, because arithmetic throughput has grown faster than bandwidth for years. That leaves a lot of work on the sloped part of the roof: elementwise ops, normalization, reductions, and LLM decode at small batch sizes. It does not leave everything there. A large, well-tiled matrix multiply clears the ridge comfortably, which is why training and prompt prefill are usually compute-bound. The next section shows how to tell from the spec sheet which side a given kernel lands on.

building the roof from a spec sheet

You do not need a profiler to draw a roofline. Two rows of the spec sheet are enough, and working it through by hand once makes the chart a profiler draws much easier to trust. The running example is an AMD MI300X.

1. Pick the right two numbers. Peak memory bandwidth is one number per card. Peak arithmetic is not: there is a separate figure for each precision, and for matrix units versus ordinary vector ALUs. Use the one your kernel actually runs on, and the dense figure rather than the sparse one, which assumes a structured-sparsity pattern your data almost certainly does not have.

MI300X roofPeak arithmeticBandwidthRidge point
BF16 on the matrix cores, dense1307 TFLOPS5.3 TB/s1307 ÷ 5.3 ≈ 247 FLOP/byte
FP32 on the vector ALUs163 TFLOPS5.3 TB/s163 ÷ 5.3 ≈ 31 FLOP/byte

2. Write the roof down. Attainable performance at intensity I is min(peak FLOPS, bandwidth × I). Evaluating it at a few points is all the chart is:

Intensity (FLOP/byte)5.3 TB/s × IAttainable, BF16 roofLimited by
15.3 TFLOPS5.3 TFLOPSMemory
1053 TFLOPS53 TFLOPSMemory
100530 TFLOPS530 TFLOPSMemory
2471307 TFLOPS1307 TFLOPSBoth: the ridge
10005300 TFLOPS1307 TFLOPSCompute

3. One roof per precision. Every roof on a card shares the same sloped part, because they all draw on the same memory. Only the flat part moves, and the ridge moves with it. On an MI300X, FP32 vector code turns compute-bound at an intensity of about 31, eight times sooner than BF16 matrix code. Narrowing the precision also raises intensity, since each value costs fewer bytes, so it moves the kernel and the roof at the same time.

Roofline for an MI300X Log-log chart of attainable TFLOPS against arithmetic intensity. A sloped memory roof at 5.3 TB/s rises until it meets a flat BF16 matrix roof at 1307 TFLOPS at an intensity of about 247. A lower dashed flat roof marks FP32 vector peak at 163 TFLOPS, meeting the slope at about 31. Four kernels are plotted: vector add and decode at batch 1 low on the slope, decode at batch 64 higher on the slope, and a 4096 matmul on the flat BF16 roof. 0.1 1 10 100 1000 10000 1 10 100 1000 10000 Arithmetic intensity, FLOP/byte (log scale) Attainable TFLOPS (log) ridge ≈ 247 FP32 vector peak, 163 TFLOPS BF16 matrix peak, 1307 TFLOPS memory roof: 5.3 TB/s × intensity vector add decode, batch 1 decode, batch 64 matmul, N = 4096
The MI300X roofs from the tables above, with the four kernels from the next table placed on them. Everything left of the ridge is limited by memory; everything right of it by the matrix cores. The same matmul in FP32 on the vector ALUs would hit the lower dashed roof instead.

4. Place the kernel. Count FLOPs and bytes for the work, assuming each operand crosses the memory bus once, and read off where it lands. The four kernels on the chart, all in BF16:

KernelFLOPsBytesIntensityCeiling on an MI300X
Vector add, n elementsn6n: two 2-byte reads, one 2-byte write≈ 0.17≈ 0.9 TFLOPS, memory-bound
LLM decode, batch 1, N×N weights2N²: a multiply and an add per weight≈ 2N²: every weight read once≈ 1≈ 5.3 TFLOPS, under 1% of peak
The same layer, batch of B2BN²≈ 2N²: weights still read once≈ BB = 64 gives ≈ 340 TFLOPS; crosses the ridge near B ≈ 250
Square matmul, N×N×N2N³6N²: A and B read, C writtenN/3Compute-bound past N ≈ 740; N = 4096 sits at ≈ 1365

The decode rows are the clearest picture of why batching matters in LLM serving: the weights cost the same bytes whether one token or a hundred uses them, so each extra token in the batch is close to free until the batch reaches the ridge. The matmul row is the ideal case. It assumes perfect reuse, which only a well-tiled kernel gets near; a naive one re-reads its operands and lands far to the left.

5. Redraw it with measured ceilings. The spec sheet gives you the ceiling in theory. For a chart you can plan against, measure both numbers on your own card: a large streaming copy (BabelStream is the usual benchmark) for bandwidth, and a large GEMM from rocBLAS or hipBLASLt (cuBLAS on Nvidia) for arithmetic. Both usually come in noticeably under the printed figures, which lowers both roofs and shifts the ridge point.

reading it

Place your kernel at its measured intensity, at its measured performance. Three cases follow.

On the sloped roof. You are bandwidth-limited and achieving close to what bandwidth allows. Adding compute, matrix units or faster math does nothing. To go faster you must move fewer bytes: fuse, improve reuse, or narrow the precision.

On the flat roof. You are compute-limited and near peak. Reducing memory traffic buys nothing. Improve the arithmetic: reach the matrix units, narrow the precision, schedule better.

Well below both. The common and least satisfying case. Neither resource is saturated, so something else is the problem: not enough parallelism in flight, too many synchronization points, divergence, or a latency chain you have not spotted. The roofline tells you to stop looking at bandwidth and arithmetic entirely.

what it is good for, and what it is not

It is good for deciding what not to try, for sanity-checking a claimed speedup, and for setting a realistic target before you start. If a kernel sits at 90% of its bandwidth roof, the remaining 10% is the whole prize, and knowing that stops you spending a week on it.

It is a simplification. It assumes traffic goes to device memory, so caches make real kernels beat the naive prediction. It ignores latency, occupancy and instruction mix. It treats a kernel as one homogeneous thing when phases may behave differently. Treat it as a bound and a direction, not a prediction.

see it on your own machine

Nsight Compute draws the chart for you and places the kernel on it, which is by far the easiest way in.

ncu --set full ./my_app | https://docs.nvidia.com/nsight-compute/ | includes a roofline chart with the kernel plotted, plus achieved bandwidth and arithmetic throughput |'rf_ncu'
ncu --metrics sm__sass_thread_inst_executed_op_fadd_pred_on.sum,dram__bytes.sum ./my_app | https://docs.nvidia.com/nsight-compute/ | the two raw ingredients, if you want to compute the ratio yourself |'rf_metrics'
rocprofv3 --kernel-trace --stats -- ./my_app | https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/ | timings on AMD; compute intensity analytically from your algorithm and divide by the measured time |'rf_rocprof'

related topics

Anatomy of a GPU — the die and the memory whose two ceilings make up the roof.
FLOPs and TFLOPS — what the peak-arithmetic figure on the spec sheet counts.
The Optimization Mindset — the informal version of this question.
The GPU Memory Hierarchy — why bytes are the scarce resource.
Tiling and Blocking for Reuse — the main technique for raising intensity.

reference

NVIDIA Nsight Compute
CUDA C++ Best Practices Guide
AMD rocprofiler