the short version

Using fewer bits per number helps twice. The matrix units perform more operations per cycle at narrower precision, and every tensor takes less bandwidth to move. Since most kernels are limited by one or the other, halving the width tends to buy something no matter which regime you are in.

The catch is range and precision, and the whole craft is knowing which parts of a computation can tolerate narrowing and which cannot.

the formats

FormatExponent / mantissaWhat it costs you
FP328 / 23Nothing, but it is the slowest and largest option
TF328 / 10Mantissa only. Same range as FP32, drop-in for many matmuls
FP165 / 10Range. Overflows and underflows easily, needs loss scaling
BF168 / 7Mantissa. Keeps FP32's full range, which is why it won
FP8 (e4m3)4 / 3Both, heavily. Needs per-tensor scaling to be usable
FP8 (e5m2)5 / 2More range, less precision. Typically used for gradients

Notice that FP16 and BF16 are the same width and differ only in where the bits go. That single design choice is the reason one of them mostly replaced the other.

why BF16 won

FP16's five exponent bits give it a narrow dynamic range. In training, gradients are frequently small enough to underflow to zero in FP16, which silently stops learning in parts of the network. The historical workaround is loss scaling: multiply the loss by a large constant before the backward pass so gradients land inside the representable range, then divide it out before the optimizer step. It works, and it is one more thing to tune and to get wrong.

BF16 simply truncates FP32's mantissa and keeps all eight exponent bits, so anything representable in FP32 is representable in BF16, just more coarsely. Underflow stops being a problem and loss scaling becomes unnecessary. Neural network training turns out to be far more tolerant of coarse mantissas than of missing range, which is why BF16 is the default on any hardware that supports it.

mixed, not low

The word "mixed" is doing real work. You do not convert everything. The standard arrangement keeps a master copy of the weights in FP32, performs the matrix multiplies with narrow inputs, and accumulates in FP32 inside the matrix unit.

That accumulator is what makes the whole thing viable. A long dot product summed in FP16 loses small contributions to rounding as the running total grows; summed in FP32 it does not. The hardware is built this way deliberately, so narrow inputs with a wide accumulator is the supported path rather than a compromise.

In practice the framework arranges this for you. Automatic mixed precision keeps a list of which operations are safe to run narrow, matmuls and convolutions, and which are not, reductions, softmax, normalization statistics and anything with a large dynamic range.

it helps memory-bound kernels too

The discussion usually centres on matrix unit throughput, which only helps if you were compute-bound. The bandwidth effect is at least as valuable and applies everywhere: a BF16 tensor is half the bytes of an FP32 one, so every elementwise kernel that touches it moves half as much data and runs roughly twice as fast.

For inference this is often the entire reason to quantize. Token generation is limited by reading weights out of memory, so halving the weight size nearly halves the time, with the arithmetic throughput being almost beside the point.

what to watch

Reductions. Summing many values narrow loses accuracy fast. Accumulate wide even when the inputs are narrow.

Anything exponential. Softmax and similar functions can overflow at narrow range, which is why the max-subtraction trick matters more, not less, in low precision.

Comparing against an FP32 baseline. Small differences are expected and are not necessarily bugs. Decide in advance what tolerance counts as correct, otherwise you cannot tell a real defect from ordinary rounding.

FP8 needs scaling infrastructure. It is not a drop-in the way BF16 is; per-tensor scale factors have to be tracked and updated, which is why it arrived first in libraries and frameworks rather than in hand-written kernels.

related topics

Tensor Cores vs Matrix Cores — the units that make narrow precision fast, and what they accumulate in.
The Optimization Mindset — which of the two benefits actually applies to your kernel.
Quantization — going below 8 bits for inference, and what it costs.

reference

NVIDIA mixed precision training guide
NVIDIA CUDA C++ Programming Guide
AMD HIP documentation