the short version

Peak TFLOPS is not a measurement. It is a multiplication: how many arithmetic lanes the chip has, times how many operations each can retire per cycle, times the clock. No kernel was run to produce it, and no kernel will ever quite reach it.

That does not make it useless. It makes it a ceiling, and knowing how the ceiling was computed tells you how much of it you can realistically expect to touch.

where the number comes from

The calculation is simple enough to do yourself: lanes × 2 × clock. The factor of two is there because a fused multiply-add counts as two floating point operations, and the hardware performs one per lane per cycle. A part with 16,384 FP32 lanes at 2 GHz therefore advertises about 65 TFLOPS of FP32, and that figure is arithmetically correct in the same way that a car's top speed is correct.

For matrix units the same logic applies with much larger numbers, because one Tensor Core or Matrix Core instruction performs many multiply-accumulates at once. This is why the headline figure on a modern datacenter part is an order of magnitude above what the general lanes could ever do, and why the two numbers are not comparable to each other.

Which clock goes into the formula. A GPU does not run at one fixed frequency. It adjusts its clock many times a second according to how much power it is drawing, how hot it is, and how much current the board can deliver. Spec sheets therefore list two figures. The base clock is the frequency the vendor guarantees the chip can hold under a sustained full load within its rated power. The boost clock is the higher figure the chip reaches when it has power and thermal headroom to spare, and it is the one peak TFLOPS is computed with.

What "boost" promises depends on the market. On datacenter parts it is effectively the ceiling: an H100 SXM tops out at 1.98 GHz and an MI300X at 2.1 GHz, which AMD calls the peak engine clock. On Nvidia's consumer cards the rated boost clock is closer to a typical value under load, and a cool card with power to spare runs above it, so a consumer card can beat its own spec sheet. The sample further down catches one doing exactly that.

The gap matters because the clock is a straight multiplier. The example part above at 2 GHz gives 65 TFLOPS; if it holds only 1.7 GHz under your workload, the same formula gives about 56 TFLOPS, 15% less before your code has done anything wrong. How close you stay to boost depends on the cooling, the power limit the card is configured with, and, ironically, on how hard your kernel works the arithmetic units, since that is what draws the most power. You can watch the live clock next to the maximum with the commands at the end of this page.

read the footnotes

Four things on a spec sheet change the number dramatically, and all four are usually in smaller type.

The footnoteWhat it does to the number
Which precisionFP64, FP32, TF32, BF16, FP16, FP8 and INT8 all get different figures, each roughly double the one above it. The headline is almost always the narrowest one.
With sparsityDoubles the figure, and requires your weights to be in a structured 2:4 sparse pattern. Dense code gets half of what is advertised.
Boost clockThe calculation uses the boost clock, not the base clock or whatever the card actually sustains under load in a warm rack.
Tensor or general lanesThe large number is the matrix units. A kernel that does not reach them is bounded by the much smaller general-lane figure.

Comparing two cards means checking that all four footnotes match before comparing the numbers, and they very often do not.

How one headline figure shrinks, H100 SXM Horizontal bars for an Nvidia H100 SXM. The headline BF16 figure with 2:4 sparsity is 1979 TFLOPS. Dense BF16 at the 1.98 GHz boost clock is 989. The same dense figure if the clock holds only 1.7 GHz is about 849. FP32 on the general lanes, without Tensor Cores, is 67. With 2:4 sparsity, the headline BF16, sparse, boost clock With 2:4 sparsity, the headline: 1,979 TFLOPS 1,979 TFLOPS Dense BF16, dense, boost clock 1.98 GHz Dense: 989 TFLOPS 989 TFLOPS Dense, clock held at 1.7 GHz 989 × 1.7 ÷ 1.98 Dense, clock held at 1.7 GHz: 849 TFLOPS 849 TFLOPS General lanes, FP32 no Tensor Cores General lanes, FP32: 67 TFLOPS 67 TFLOPS
Three of the four footnotes applied to one card. Only the top bar is printed in large type; your code runs somewhere at or below the bottom three. The 1.7 GHz row is an illustration of a clock held below boost under load, not a measured figure.

the four reasons you fall short

Even with matching footnotes, a real kernel lands below the ceiling. The first reason decides which ceiling you are under at all; the other three are why you fall short of it.

1. You may be memory-bound, not compute-bound. This is the big one. Every GPU has two ceilings: how fast it can do arithmetic (peak FLOPS) and how fast it can move data to and from device memory (peak bandwidth, in TB/s). Which one limits a kernel depends on its arithmetic intensity, the FLOPs it performs for every byte it moves. Divide the card's peak FLOPS by its bandwidth and you get the ridge point, the intensity a kernel needs before arithmetic becomes the limit. The best a kernel can reach is min(peak FLOPS, bandwidth × intensity).

An AMD MI300X, for example, does about 1307 TFLOPS of dense BF16 on its matrix cores against 5.3 TB/s of bandwidth, a ridge point of roughly 247 FLOPs per byte. A vector add does one FLOP per 6 bytes of BF16 traffic, so its ceiling is about 5.3 × 0.17 ≈ 0.9 TFLOPS: well under 1% of peak, and nothing inside the kernel can fix that. A large, well-tiled matrix multiply clears the ridge easily and is compute-bound, which is the only case where the peak FLOPS figure is the right yardstick. Elementwise ops, normalization, reductions and LLM decode at small batch sizes sit on the memory side; measure them in bytes per second against peak bandwidth instead. The Roofline Model works this through in full.

Once you know you are compute-bound, three more things keep you below the flat ceiling.

2. Not every lane is working. Divergent branches mask lanes off, and a block size that is not a multiple of the group size leaves a tail permanently idle.

3. Not every instruction is arithmetic. Address calculation, loop control, loads and stores all occupy issue slots that are then not issuing multiply-adds.

4. The clock is not the boost clock. As described above, peak assumes the boost clock. Sustained heavy arithmetic is exactly what pushes a card into its power and thermal limits, so on a datacenter part the clock during your kernel usually sits somewhere between base and boost, and every percent it drops comes straight off the ceiling. The closer you get to compute-bound, the more this bites.

what fraction is actually good

Kind of kernelRealistic share of peak
A carefully tuned large dense matrix multiplyHigh. Vendor BLAS gets a large majority of peak, and this is the case the hardware was designed around.
A decent first attempt at a tiled matmulA modest fraction. Getting from here to the line above is most of what kernel optimization is.
An elementwise or normalization kernelA tiny fraction, and that is correct. It is limited by bandwidth, so FLOPs is the wrong yardstick entirely.
Single-request LLM token generationVery low, for the same reason. Measure bytes per second instead.

The useful habit is to decide which row you are in before optimizing. Chasing FLOPs in a bandwidth-limited kernel is the most common way to spend a week and gain nothing.

measure yours instead

Achieved FLOPs is a division you can do from two numbers you already have: count the floating point operations your algorithm performs, which is usually a simple formula, and divide by the measured runtime. For a matrix multiply of size N it is 2N³ operations, and that one formula covers most of what matters in deep learning.

Comparing that against peak gives you a percentage that means something, and comparing it against the vendor library on the same problem gives you a target. A profiler will also report it directly, along with whether you were bounded by arithmetic or by memory.

try it

The companion sample measures both halves of this page on your own GPU: achieved FP32 throughput, from threads running long chains of independent fused multiply-adds with no memory traffic, and the clock the GPU held while doing it, from each block counting its own cycles. Pass the spec-sheet figure and it prints the percentage:


hipcc -O3 -o main flops_achieved_vs_peak.hip
./main 163.4        # MI300X FP32 vector peak, for example
            

On a small Nvidia card, built as CUDA, it measured 1.23 TFLOPS against a spec-sheet 1.13, or 109% of peak, at a clock of about 1.64 GHz. The spec sheet computes its figure at a rated boost clock of 1.47 GHz. The card was cool and had power to spare, so it boosted past that, and 384 lanes × 2 × 1.64 GHz ≈ 1.26 TFLOPS: the lanes were almost fully busy, and the clock alone explains the rest. Run the same program on a datacenter part under sustained load and expect the opposite, a clock below boost and a result below peak. Either way, the clock is the number that moved.

The full program, with build instructions for hipcc and notes for nvcc: TopNotchNote/gpu/flops_achieved_vs_peak.hip

see it on your own machine

nvidia-smi --query-gpu=name,clocks.max.sm,clocks.sm,power.draw,power.limit --format=csv | https://developer.nvidia.com/nvidia-system-management-interface | the boost clock next to the clock right now, and the power headroom that decides which one you get |'ft_smi'
ncu --set full ./my_app | https://docs.nvidia.com/nsight-compute/ | per-kernel achieved FLOPs, achieved bandwidth, and which of the two was your limiter |'ft_ncu'
rocm-smi --showclocks --showpower | https://rocm.docs.amd.com/projects/rocm_smi_lib/en/latest/ | sustained clocks and board power on AMD hardware, the same sanity check |'ft_rocmsmi'

rules of thumb

  • Before comparing two cards, match all four footnotes: precision, sparsity, clock, and matrix versus general lanes.
  • Work out whether your kernel is compute-bound or memory-bound before you judge it by FLOPs at all.
  • Judge a compute-bound kernel against the vendor library on the same problem, not against the spec sheet.
  • Watch the clock while you benchmark. A result that moves between runs is often the clock moving, not your code.
  • Report achieved TFLOPS with the precision, problem size and card, or the number means nothing to anyone else.

related topics

Tensor Cores vs Matrix Cores — where the very large numbers on the spec sheet come from.
The Roofline Model — the compute-bound versus memory-bound check, drawn as a chart.
The GPU Memory Hierarchy — where the bytes a memory-bound kernel waits on come from.
SM vs CU — the lane counts that go into the formula, for two real chips.
Anatomy of a GPU — reading the rest of a spec sheet without being fooled.

reference

NVIDIA CUDA C++ Programming Guide
NVIDIA GPU architecture whitepapers
AMD CDNA architecture
NVIDIA Nsight Compute