Tensor Cores vs Matrix Cores
A separate datapath that multiplies small matrices in one instruction, and supplies most of the FLOPs on the chip.
Alongside the ordinary arithmetic lanes, modern GPUs carry a second kind of unit that does one thing: multiply two small matrices and add the result to a third, as a single instruction. Nvidia calls them Tensor Cores, AMD calls them Matrix Cores and drives them with MFMA instructions. On any current datacenter part they supply the large majority of the chip's advertised FLOPs.
If you are running matrix multiplication and not reaching these units, you are using a small fraction of the hardware you paid for.
The operation is a fused matrix multiply-accumulate on fixed-size tiles: D = A × B + C, where A, B, C and D are small matrices held across the registers of a whole lane group rather than a single thread. One instruction consumes the entire tile and produces the entire result.
The tile shapes are fixed by the hardware, something like 16×16×16 depending on architecture and precision, and the operands are not laid out the way you would naturally store a matrix. Each lane holds a specific, architecture-defined fragment of the tile. Getting data into that arrangement is most of the work of using these units, and is the reason the programming model looks nothing like ordinary GPU code.
Two reasons, and neither is magic. First, operand reuse: a general lane performing a dot product fetches operands, multiplies, accumulates, and repeats, paying register file traffic on every step. A matrix unit reads a tile once and performs every multiply-accumulate in that tile inside a fixed-function datapath, so the same operands feed many more operations. Second, precision: these units are built for narrow inputs, and narrow inputs mean more multipliers fit in the same silicon.
The combination is worth roughly an order of magnitude of arithmetic throughput over the general lanes at the same precision. It is also why the headline TFLOPs on a spec sheet is almost always a matrix-unit number at a low precision, not something the ordinary FP32 lanes could ever reach.
| Input precision | Typically accumulates in | Used for |
|---|---|---|
| FP16 | FP32 | Training and inference, the long-standing default |
| BF16 | FP32 | Training, preferred for its wider exponent range |
| TF32 | FP32 | A drop-in for FP32 matmuls with reduced mantissa |
| FP8 | FP16 or FP32 | Inference, and increasingly training on recent parts |
| INT8 / INT4 | INT32 | Quantized inference |
The accumulate column is the part people skip and should not. Inputs are narrow but the running sum is kept wide, which is what makes these units usable for real numerical work rather than a curiosity. BF16 is generally preferred over FP16 for training because it keeps FP32's exponent range and sacrifices mantissa bits instead, which makes overflow far less likely and loss scaling largely unnecessary.
| Nvidia | AMD CDNA | |
|---|---|---|
| Name | Tensor Cores | Matrix Cores |
| Instruction family | mma / wmma, and asynchronous variants on recent parts | MFMA |
| Intrinsic-level API | The wmma interface in CUDA C++ | MFMA builtins in HIP |
| Library that does it for you | cuBLAS, cuDNN | rocBLAS, hipBLASLt, MIOpen |
| Template layer for custom kernels | CUTLASS with CuTe | Composable Kernel |
You can write to these units by hand, and for a while everyone who wanted peak performance had to. It is unpleasant work: the fragment layouts are architecture-specific, the scratchpad needs swizzling to avoid bank conflicts when feeding them, and the result is fast on exactly the hardware you tuned it for.
For ordinary matrix multiplication the vendor libraries already do this, are tuned per architecture, and will beat a first hand-written attempt comfortably. The reason to go lower is fusion: when you need the matrix multiply to be part of a larger operation so that intermediate results never leave the chip, no library call expresses that, and you reach for a template layer like CUTLASS or Composable Kernel, or a compiler like Triton. That is precisely the situation a fused attention kernel is in.
These units accelerate arithmetic, so they help exactly when arithmetic is the constraint. A kernel limited by memory bandwidth gains nothing from them, because the lanes were already idle waiting on data and making them faster idles them harder.
This is why a large batched training step, which is dense matrix work, sees enormous benefit, while single-request token generation sees very little: the latter is moving weights, not multiplying them. Knowing which regime a kernel is in before reaching for lower precision is the difference between a real speedup and an afternoon spent introducing numerical risk for nothing.
Whether the units exist at all depends on the part, and whether your kernel reached them is a profiler question.