Mixed Precision: FP16, BF16 and FP8
Narrower inputs buy you arithmetic throughput and bandwidth at once. The accumulator is what keeps it correct.
Using fewer bits per number helps twice. The matrix units perform more operations per cycle at narrower precision, and every tensor takes less bandwidth to move. Since most kernels are limited by one or the other, halving the width tends to buy something no matter which regime you are in.
The catch is range and precision, and the whole craft is knowing which parts of a computation can tolerate narrowing and which cannot.
| Format | Exponent / mantissa | What it costs you |
|---|---|---|
| FP32 | 8 / 23 | Nothing, but it is the slowest and largest option |
| TF32 | 8 / 10 | Mantissa only. Same range as FP32, drop-in for many matmuls |
| FP16 | 5 / 10 | Range. Overflows and underflows easily, needs loss scaling |
| BF16 | 8 / 7 | Mantissa. Keeps FP32's full range, which is why it won |
| FP8 (e4m3) | 4 / 3 | Both, heavily. Needs per-tensor scaling to be usable |
| FP8 (e5m2) | 5 / 2 | More range, less precision. Typically used for gradients |
Notice that FP16 and BF16 are the same width and differ only in where the bits go. That single design choice is the reason one of them mostly replaced the other.
FP16's five exponent bits give it a narrow dynamic range. In training, gradients are frequently small enough to underflow to zero in FP16, which silently stops learning in parts of the network. The historical workaround is loss scaling: multiply the loss by a large constant before the backward pass so gradients land inside the representable range, then divide it out before the optimizer step. It works, and it is one more thing to tune and to get wrong.
BF16 simply truncates FP32's mantissa and keeps all eight exponent bits, so anything representable in FP32 is representable in BF16, just more coarsely. Underflow stops being a problem and loss scaling becomes unnecessary. Neural network training turns out to be far more tolerant of coarse mantissas than of missing range, which is why BF16 is the default on any hardware that supports it.
The word "mixed" is doing real work. You do not convert everything. The standard arrangement keeps a master copy of the weights in FP32, performs the matrix multiplies with narrow inputs, and accumulates in FP32 inside the matrix unit.
That accumulator is what makes the whole thing viable. A long dot product summed in FP16 loses small contributions to rounding as the running total grows; summed in FP32 it does not. The hardware is built this way deliberately, so narrow inputs with a wide accumulator is the supported path rather than a compromise.
In practice the framework arranges this for you. Automatic mixed precision keeps a list of which operations are safe to run narrow, matmuls and convolutions, and which are not, reductions, softmax, normalization statistics and anything with a large dynamic range.
The discussion usually centres on matrix unit throughput, which only helps if you were compute-bound. The bandwidth effect is at least as valuable and applies everywhere: a BF16 tensor is half the bytes of an FP32 one, so every elementwise kernel that touches it moves half as much data and runs roughly twice as fast.
For inference this is often the entire reason to quantize. Token generation is limited by reading weights out of memory, so halving the weight size nearly halves the time, with the arithmetic throughput being almost beside the point.
Reductions. Summing many values narrow loses accuracy fast. Accumulate wide even when the inputs are narrow.
Anything exponential. Softmax and similar functions can overflow at narrow range, which is why the max-subtraction trick matters more, not less, in low precision.
Comparing against an FP32 baseline. Small differences are expected and are not necessarily bugs. Decide in advance what tolerance counts as correct, otherwise you cannot tell a real defect from ordinary rounding.
FP8 needs scaling infrastructure. It is not a drop-in the way BF16 is; per-tensor scale factors have to be tracked and updated, which is why it arrived first in libraries and frameworks rather than in hand-written kernels.