FLOPs and TFLOPS: Why the Spec Number Is Not What You Get
Where the headline figure comes from, and the four reasons your kernel never reaches it.
Peak TFLOPS is not a measurement. It is a multiplication: how many arithmetic lanes the chip has, times how many operations each can retire per cycle, times the clock. No kernel was run to produce it, and no kernel will ever quite reach it.
That does not make it useless. It makes it a ceiling, and knowing how the ceiling was computed tells you how much of it you can realistically expect to touch.
The calculation is simple enough to do yourself: lanes × 2 × clock. The factor of two is there because a fused multiply-add counts as two floating point operations, and the hardware performs one per lane per cycle. A part with 16,384 FP32 lanes at 2 GHz therefore advertises about 65 TFLOPS of FP32, and that figure is arithmetically correct in the same way that a car's top speed is correct.
For matrix units the same logic applies with much larger numbers, because one Tensor Core or Matrix Core instruction performs many multiply-accumulates at once. This is why the headline figure on a modern datacenter part is an order of magnitude above what the general lanes could ever do, and why the two numbers are not comparable to each other.
Which clock goes into the formula. A GPU does not run at one fixed frequency. It adjusts its clock many times a second according to how much power it is drawing, how hot it is, and how much current the board can deliver. Spec sheets therefore list two figures. The base clock is the frequency the vendor guarantees the chip can hold under a sustained full load within its rated power. The boost clock is the higher figure the chip reaches when it has power and thermal headroom to spare, and it is the one peak TFLOPS is computed with.
What "boost" promises depends on the market. On datacenter parts it is effectively the ceiling: an H100 SXM tops out at 1.98 GHz and an MI300X at 2.1 GHz, which AMD calls the peak engine clock. On Nvidia's consumer cards the rated boost clock is closer to a typical value under load, and a cool card with power to spare runs above it, so a consumer card can beat its own spec sheet. The sample further down catches one doing exactly that.
The gap matters because the clock is a straight multiplier. The example part above at 2 GHz gives 65 TFLOPS; if it holds only 1.7 GHz under your workload, the same formula gives about 56 TFLOPS, 15% less before your code has done anything wrong. How close you stay to boost depends on the cooling, the power limit the card is configured with, and, ironically, on how hard your kernel works the arithmetic units, since that is what draws the most power. You can watch the live clock next to the maximum with the commands at the end of this page.
Four things on a spec sheet change the number dramatically, and all four are usually in smaller type.
| The footnote | What it does to the number |
|---|---|
| Which precision | FP64, FP32, TF32, BF16, FP16, FP8 and INT8 all get different figures, each roughly double the one above it. The headline is almost always the narrowest one. |
| With sparsity | Doubles the figure, and requires your weights to be in a structured 2:4 sparse pattern. Dense code gets half of what is advertised. |
| Boost clock | The calculation uses the boost clock, not the base clock or whatever the card actually sustains under load in a warm rack. |
| Tensor or general lanes | The large number is the matrix units. A kernel that does not reach them is bounded by the much smaller general-lane figure. |
Comparing two cards means checking that all four footnotes match before comparing the numbers, and they very often do not.
Even with matching footnotes, a real kernel lands below the ceiling. The first reason decides which ceiling you are under at all; the other three are why you fall short of it.
1. You may be memory-bound, not compute-bound. This is the big one. Every GPU has two ceilings: how fast it can do arithmetic (peak FLOPS) and how fast it can move data to and from device memory (peak bandwidth, in TB/s). Which one limits a kernel depends on its arithmetic intensity, the FLOPs it performs for every byte it moves. Divide the card's peak FLOPS by its bandwidth and you get the ridge point, the intensity a kernel needs before arithmetic becomes the limit. The best a kernel can reach is min(peak FLOPS, bandwidth × intensity).
An AMD MI300X, for example, does about 1307 TFLOPS of dense BF16 on its matrix cores against 5.3 TB/s of bandwidth, a ridge point of roughly 247 FLOPs per byte. A vector add does one FLOP per 6 bytes of BF16 traffic, so its ceiling is about 5.3 × 0.17 ≈ 0.9 TFLOPS: well under 1% of peak, and nothing inside the kernel can fix that. A large, well-tiled matrix multiply clears the ridge easily and is compute-bound, which is the only case where the peak FLOPS figure is the right yardstick. Elementwise ops, normalization, reductions and LLM decode at small batch sizes sit on the memory side; measure them in bytes per second against peak bandwidth instead. The Roofline Model works this through in full.
Once you know you are compute-bound, three more things keep you below the flat ceiling.
2. Not every lane is working. Divergent branches mask lanes off, and a block size that is not a multiple of the group size leaves a tail permanently idle.
3. Not every instruction is arithmetic. Address calculation, loop control, loads and stores all occupy issue slots that are then not issuing multiply-adds.
4. The clock is not the boost clock. As described above, peak assumes the boost clock. Sustained heavy arithmetic is exactly what pushes a card into its power and thermal limits, so on a datacenter part the clock during your kernel usually sits somewhere between base and boost, and every percent it drops comes straight off the ceiling. The closer you get to compute-bound, the more this bites.
| Kind of kernel | Realistic share of peak |
|---|---|
| A carefully tuned large dense matrix multiply | High. Vendor BLAS gets a large majority of peak, and this is the case the hardware was designed around. |
| A decent first attempt at a tiled matmul | A modest fraction. Getting from here to the line above is most of what kernel optimization is. |
| An elementwise or normalization kernel | A tiny fraction, and that is correct. It is limited by bandwidth, so FLOPs is the wrong yardstick entirely. |
| Single-request LLM token generation | Very low, for the same reason. Measure bytes per second instead. |
The useful habit is to decide which row you are in before optimizing. Chasing FLOPs in a bandwidth-limited kernel is the most common way to spend a week and gain nothing.
Achieved FLOPs is a division you can do from two numbers you already have: count the floating point operations your algorithm performs, which is usually a simple formula, and divide by the measured runtime. For a matrix multiply of size N it is 2N³ operations, and that one formula covers most of what matters in deep learning.
Comparing that against peak gives you a percentage that means something, and comparing it against the vendor library on the same problem gives you a target. A profiler will also report it directly, along with whether you were bounded by arithmetic or by memory.
The companion sample measures both halves of this page on your own GPU: achieved FP32 throughput, from threads running long chains of independent fused multiply-adds with no memory traffic, and the clock the GPU held while doing it, from each block counting its own cycles. Pass the spec-sheet figure and it prints the percentage:
hipcc -O3 -o main flops_achieved_vs_peak.hip
./main 163.4 # MI300X FP32 vector peak, for example
On a small Nvidia card, built as CUDA, it measured 1.23 TFLOPS against a spec-sheet 1.13, or 109% of peak, at a clock of about 1.64 GHz. The spec sheet computes its figure at a rated boost clock of 1.47 GHz. The card was cool and had power to spare, so it boosted past that, and 384 lanes × 2 × 1.64 GHz ≈ 1.26 TFLOPS: the lanes were almost fully busy, and the clock alone explains the rest. Run the same program on a datacenter part under sustained load and expect the opposite, a clock below boost and a result below peak. Either way, the clock is the number that moved.
hipcc and notes for nvcc:
TopNotchNote/gpu/flops_achieved_vs_peak.hip