the short version

Alongside the ordinary arithmetic lanes, modern GPUs carry a second kind of unit that does one thing: multiply two small matrices and add the result to a third, as a single instruction. Nvidia calls them Tensor Cores, AMD calls them Matrix Cores and drives them with MFMA instructions. On any current datacenter part they supply the large majority of the chip's advertised FLOPs.

If you are running matrix multiplication and not reaching these units, you are using a small fraction of the hardware you paid for.

what they actually compute

The operation is a fused matrix multiply-accumulate on fixed-size tiles: D = A × B + C, where A, B, C and D are small matrices held across the registers of a whole lane group rather than a single thread. One instruction consumes the entire tile and produces the entire result.

The tile shapes are fixed by the hardware, something like 16×16×16 depending on architecture and precision, and the operands are not laid out the way you would naturally store a matrix. Each lane holds a specific, architecture-defined fragment of the tile. Getting data into that arrangement is most of the work of using these units, and is the reason the programming model looks nothing like ordinary GPU code.

why they are so much faster

Two reasons, and neither is magic. First, operand reuse: a general lane performing a dot product fetches operands, multiplies, accumulates, and repeats, paying register file traffic on every step. A matrix unit reads a tile once and performs every multiply-accumulate in that tile inside a fixed-function datapath, so the same operands feed many more operations. Second, precision: these units are built for narrow inputs, and narrow inputs mean more multipliers fit in the same silicon.

The combination is worth roughly an order of magnitude of arithmetic throughput over the general lanes at the same precision. It is also why the headline TFLOPs on a spec sheet is almost always a matrix-unit number at a low precision, not something the ordinary FP32 lanes could ever reach.

precision, and what accumulates where

Input precisionTypically accumulates inUsed for
FP16FP32Training and inference, the long-standing default
BF16FP32Training, preferred for its wider exponent range
TF32FP32A drop-in for FP32 matmuls with reduced mantissa
FP8FP16 or FP32Inference, and increasingly training on recent parts
INT8 / INT4INT32Quantized inference

The accumulate column is the part people skip and should not. Inputs are narrow but the running sum is kept wide, which is what makes these units usable for real numerical work rather than a curiosity. BF16 is generally preferred over FP16 for training because it keeps FP32's exponent range and sacrifices mantissa bits instead, which makes overflow far less likely and loss scaling largely unnecessary.

the two, side by side

NvidiaAMD CDNA
NameTensor CoresMatrix Cores
Instruction familymma / wmma, and asynchronous variants on recent partsMFMA
Intrinsic-level APIThe wmma interface in CUDA C++MFMA builtins in HIP
Library that does it for youcuBLAS, cuDNNrocBLAS, hipBLASLt, MIOpen
Template layer for custom kernelsCUTLASS with CuTeComposable Kernel

why you probably should not program them directly

You can write to these units by hand, and for a while everyone who wanted peak performance had to. It is unpleasant work: the fragment layouts are architecture-specific, the scratchpad needs swizzling to avoid bank conflicts when feeding them, and the result is fast on exactly the hardware you tuned it for.

For ordinary matrix multiplication the vendor libraries already do this, are tuned per architecture, and will beat a first hand-written attempt comfortably. The reason to go lower is fusion: when you need the matrix multiply to be part of a larger operation so that intermediate results never leave the chip, no library call expresses that, and you reach for a template layer like CUTLASS or Composable Kernel, or a compiler like Triton. That is precisely the situation a fused attention kernel is in.

when they do not help at all

These units accelerate arithmetic, so they help exactly when arithmetic is the constraint. A kernel limited by memory bandwidth gains nothing from them, because the lanes were already idle waiting on data and making them faster idles them harder.

This is why a large batched training step, which is dense matrix work, sees enormous benefit, while single-request token generation sees very little: the latter is moving weights, not multiplying them. Knowing which regime a kernel is in before reaching for lower precision is the difference between a real speedup and an afternoon spent introducing numerical risk for nothing.

see it on your own machine

Whether the units exist at all depends on the part, and whether your kernel reached them is a profiler question.

nvidia-smi --query-gpu=name,compute_cap --format=csv | https://developer.nvidia.com/nvidia-system-management-interface | compute capability tells you which Tensor Core generation and which precisions are available |'tc_smi'
rocminfo | https://rocm.docs.amd.com/projects/rocminfo/en/latest/ | the gfx target identifies whether the part carries Matrix Cores and which MFMA variants it supports |'tc_rocminfo'
ncu --metrics sm__inst_executed_pipe_tensor.sum ./my_app | https://docs.nvidia.com/nsight-compute/ | counts instructions actually issued to the Tensor Core pipe; zero means you never reached them |'tc_ncu'

related topics

SM vs CU — the compute block these units live inside.
GPU Libraries — cuBLAS, rocBLAS and the rest, which drive these units for you.
The GPU Memory Hierarchy — why feeding these units is usually the hard part.

reference

CUDA C++ Programming Guide: warp matrix functions
AMD HIP documentation
NVIDIA CUTLASS
AMD Composable Kernel