three things bolted together

What people call "a GPU" is really an assembly of three parts, and it pays to keep them separate in your head because each one imposes a different ceiling:

PartWhat it isWhat it limits
The dieSilicon covered in SMs (Nvidia) or CUs (AMD)Peak arithmetic — your FLOPs ceiling
The memoryHBM stacks on the package, or GDDR chips on the boardBandwidth and capacity — how fast and how much
The host linkA PCIe slot, plus NVLink or Infinity Fabric on datacenter partsHow fast data gets on and off the card at all
Top-down view of a GPU package One GPU die on a package substrate, flanked by eight HBM stacks joined to it by wide, short buses. Below the package, a much narrower PCIe link runs to the host CPU and a second link runs to peer GPUs. Package HBM HBM HBM HBM HBM HBM HBM HBM L2 cache GPU die SMs (Nvidia) / CUs (AMD) PCIe — to host CPU NVLink / Infinity Fabric 1 2 3 1 The die your FLOPs ceiling 2 The memory TB/s, on-package 3 The host link tens of GB/s
The same three parts as the table above. Note the bus widths: the path from the stacks to the die is short and very wide, the path off the package is neither.

A kernel is bound by whichever of the three runs out first, and which one that is depends on the work as much as the card. A large matrix multiply runs out of arithmetic; generating text one token at a time runs out of bandwidth. Two rows of the spec sheet are enough to tell which, and a section below shows the arithmetic.

Two words are about to do a lot of work on this page, so they are worth pinning down. A consumer card is the kind that goes in a desktop or workstation — GeForce on Nvidia, Radeon on AMD — designed for graphics first and sold at retail. A datacenter card — H100, MI300 and their successors — is designed for compute in a server rack: several times the memory, a dedicated high-speed link to its neighbours, usually no display output at all, and an order of magnitude more expensive. All three parts above are built differently depending on which of the two you are holding, and the next three sections are largely about how.

the die

The die is one piece of silicon tiled with dozens to hundreds of identical compute blocks — Streaming Multiprocessors on Nvidia, Compute Units on AMD — plus a shared L2 cache, the memory controllers, and the scheduling hardware that hands work to the blocks. The design is deliberately repetitive: the same block stamped out many times, which is what lets a vendor ship a bigger and a smaller version of the same architecture by simply including more or fewer of them, and what lets a partially defective die be sold as a lower-tier part with some blocks disabled.

Two consequences worth carrying forward. First, your work only scales if you give it enough independent pieces to occupy every block — a kernel that fills ten blocks on a hundred-block die wastes ninety percent of the chip no matter how well written it is. Second, "how many cores does it have" is a nearly meaningless question across vendors, because Nvidia and AMD count the insides of these blocks differently.

the memory, and why it is so fast

GPU memory achieves its bandwidth by being wide and close, not by being clocked absurdly high. Two packaging styles do this differently:

HBM (datacenter)GDDR (consumer)
Where it sitsStacked DRAM dies on the same package as the GPUSeparate chips soldered around the die on the board
InterfaceExtremely wide, relatively modest clocksNarrow per chip, very high clocks
Typical capacityLarge — well past 100 GB on current datacenter parts, and climbing every generationSmaller — roughly 8–32 GB
Cost and powerExpensive, power-efficient per byte movedCheap, less efficient per byte
Found onH100, MI300 and similarGeForce, Radeon
HBM versus GDDR packaging On the left, HBM: DRAM dies stacked on the same substrate as the GPU die, joined by two very wide short buses. On the right, GDDR: eight separate memory chips soldered around the die on the board, each reaching it over a long narrow trace. HBM — datacenter package substrate stacked stacked die On the same substrate as the die. Very wide bus, short traces, modest clocks. GDDR — consumer board GDDR GDDR GDDR GDDR GDDR GDDR GDDR GDDR die Separate chips soldered around the die. Narrow bus per chip, made up with very high clocks.
Two ways to reach high bandwidth. HBM buys it with width and proximity; GDDR buys it with clock speed over longer traces.

Either way the bus between die and memory is far wider and far shorter than the path from a CPU to its DIMM slots, which is the entire reason GPU memory bandwidth is measured in terabytes per second while a desktop CPU's is measured in tens of gigabytes. You pay for that with capacity: a GPU has one to two orders of magnitude less memory than the server it is plugged into, and running out of it is the single most common practical wall in deep learning.

the link to the host

From the operating system's point of view the whole card is just a PCIe device — the driver enumerates it much like an NVMe SSD or a network card, only with a much larger BAR, the memory-mapped window through which the host can reach device memory. That link is dramatically slower than the on-package memory bus: tens of gigabytes per second against thousands. Any design that copies data back and forth per iteration is bound by this number, which is why the standing advice is to move data onto the device once and keep it there.

Datacenter cards add a second, much faster network for GPU-to-GPU traffic — NVLink with NVSwitch on Nvidia, Infinity Fabric / xGMI on AMD — so that peer traffic in multi-GPU training does not have to be routed through the host at PCIe speed. Consumer cards generally do not have this, which is one of the real functional differences between a gaming card and a datacenter one.

reading a spec sheet without being fooled

Four rows do most of the work, and each has a catch:

Spec rowWhat it really tells you
Peak TFLOPSA theoretical ceiling at a specific precision. Real kernels reach a fraction of it, and the headline figure is often the sparse or low-precision number.
Memory bandwidthThe ceiling for any kernel that does little arithmetic per byte it moves: elementwise ops, normalization, LLM decode. Divide peak TFLOPS by it to find out whether yours is one of them.
Memory capacityA hard wall, not a soft one. It decides which models fit at all.
Board power (TDP)What the card is allowed to draw. Sustained clocks, and therefore sustained throughput, depend on cooling holding up.

compute-bound or memory-bound: two rows and a division

Peak TFLOPS and memory bandwidth together answer the question every optimization starts from: will this kernel run out of arithmetic first, or out of bytes? It takes three steps.

1. The card's ridge point. Divide peak FLOPS by memory bandwidth, using the precision you actually run and the dense figure, not the sparse one. For an MI300X at BF16 that is 1307 TFLOPS ÷ 5.3 TB/s ≈ 247 FLOPs per byte. It is how much arithmetic a kernel has to do for every byte it moves before memory stops being the bottleneck.

2. The kernel's arithmetic intensity. Count the FLOPs the kernel performs and the bytes it reads from and writes to device memory, and divide. This is a property of the work, not of the card.

3. Compare. The best the kernel can do is min(peak FLOPS, bandwidth × intensity). Below the ridge point the second term is the smaller one and memory is the limit. Above it, the die is.

Workload (BF16)FLOPsBytes movedIntensityOn an MI300X
Vector add, n elementsn6n≈ 0.17Memory-bound, at most ≈ 0.9 TFLOPS
LLM decode, one token through an N×N weight matrix2N²≈ 2N² (the weights)≈ 1Memory-bound, at most ≈ 5 TFLOPS: under 1% of peak
The same layer, batch of B tokens2BN²≈ 2N²≈ BCrosses the ridge around B ≈ 250
Square matmul, N×N×N2N³6N²N/3Compute-bound past N ≈ 740; at N = 4096 the intensity is ≈ 1365

So neither half of the chip is "the" bottleneck in general. Training and the prefill phase of inference are dominated by large matrix multiplies and sit on the compute side. Token-by-token decoding, elementwise ops and normalization sit on the memory side. Precision moves the line as well: the same MI300X peaks at about 163 TFLOPS for plain FP32 vector math, a ridge point near 31, so FP32 code turns compute-bound far sooner.

Treat the results as ceilings, not predictions. Real kernels reach only part of either peak, so for planning, redo the division with the bandwidth and FLOPS you actually measure. The byte counts also assume every operand crosses the memory bus exactly once. A poorly tiled matmul re-reads its inputs and lands far below N/3. Plotting the ceiling across every intensity gives the roofline chart, which The Roofline Model covers along with how to read it.

see it on your own machine

Each of the three parts is visible from the command line.

lspci -d 10de: | https://man7.org/linux/man-pages/man8/lspci.8.html | confirms the card is enumerated as a PCIe device on Nvidia hardware (10de is Nvidia's PCI vendor ID) |'anat_lspci_nv'
lspci -d 1002: | https://man7.org/linux/man-pages/man8/lspci.8.html | the same check on AMD hardware (1002 is AMD's PCI vendor ID) |'anat_lspci_amd'
nvidia-smi --query-gpu=name,memory.total,pcie.link.gen.current,pcie.link.width.current --format=csv | https://developer.nvidia.com/nvidia-system-management-interface | memory capacity plus the PCIe generation and lane width actually negotiated, which is often lower than the slot's rating |'anat_nvsmi_link'
nvidia-smi -q -d MEMORY,POWER | https://developer.nvidia.com/nvidia-system-management-interface | the full memory and power blocks, including the enforced power limit the card is throttling to |'anat_nvsmi_q'
rocm-smi --showmeminfo vram --showpower | https://rocm.docs.amd.com/projects/rocm_smi_lib/en/latest/ | VRAM usage and current board power on AMD hardware |'anat_rocmsmi'

related topics

CPU vs GPU: Why GPUs Exist — the design bet this hardware is built to win.
The Roofline Model — the ridge-point calculation above, drawn as a chart and used to pick optimizations.
FLOPs and TFLOPS — what the peak-arithmetic row on the spec sheet actually counts.
GPU Optimization — working within whichever ceiling you hit.
GPU Memory Planning for LLMs — what the capacity wall means in practice when serving a model.

reference

NVIDIA GPU architecture whitepapers
AMD CDNA architecture
JEDEC HBM standard
NVIDIA NVLink
AMD Infinity Architecture