the short version

A CPU is built to finish one task as soon as possible. A GPU is built to finish a million tasks per second, and does not care how long any single one of them takes. Both chips get roughly the same budget — a few hundred square millimetres of silicon and a power envelope — and they spend it in opposite directions. Everything else on this page, and most of what makes GPU code fast or slow, follows from that one choice.

where the transistors go

The interesting comparison is not "how many cores" but what the silicon is spent on. A CPU spends most of its area on machinery that makes a single instruction stream run fast: large caches, out-of-order scheduling, register renaming, branch predictors, deep pipelines. None of that machinery does arithmetic — it exists to keep the few arithmetic units fed. A GPU deletes almost all of it and spends the area on arithmetic units instead.

Budget spent onCPUGPU
Independent coresTens (8–128 on servers)Dozens to hundreds of SMs/CUs, each with many lanes
Arithmetic lanesHundreds to a couple thousand, via SIMD (AVX-512, SVE)Tens of thousands
Cache per laneVery large (tens of MB of L3)Very small (tens to a couple hundred KB on-chip per SM/CU)
Out-of-order executionYes, aggressiveNo — in-order issue per warp/wavefront
Branch predictionYes, sophisticatedEssentially none
Clock speedHigh (4–6 GHz boost)Lower (roughly 1–2.5 GHz)
Memory priorityLow latencyHigh bandwidth

Notice the last row especially. A CPU's memory system is tuned to return one cache line as quickly as possible; a GPU's is tuned to move enormous volumes of data per second even though any individual access is slower than the CPU's. Those are not the same goal, and you cannot maximize both.

latency hiding: the actual mechanism

Here is the part that is easy to miss. A GPU does not make memory faster — it makes memory slower. A read that misses cache and goes out to GPU memory takes on the order of 400–600 nanoseconds, against roughly 80–100 for a CPU reaching its own DRAM. Like for like, at the same level of the hierarchy, the GPU loses. It wins anyway because of what it does while waiting.

A CPU facing a slow memory access tries to avoid the stall: it predicts the branch, prefetches the line, reorders later instructions to run ahead. All of that costs transistors and power. A GPU facing the same stall simply switches to another group of threads that is ready to run, and issues their instruction this cycle instead. Each SM/CU keeps far more threads resident than it can execute at once, precisely so that there is always someone else to run.

This is why "just add more parallelism" is the standard GPU fix for a slow kernel, and why a GPU running only a handful of threads is catastrophically slow — not because the threads are slow, but because there is nobody left to hide behind. It is also the reason occupancy becomes a thing you tune later on.

what the bet predicts about your code

Once you accept the throughput bet, most GPU performance advice stops being a list to memorize and becomes derivable:

Because the GPU……your code must
hides latency with spare threadsexpose thousands of independent work items, not dozens
executes lanes in lockstep groupsavoid divergent branches inside a group
optimizes for bandwidth, not latencyread memory in wide, contiguous, neighbour-friendly patterns
has tiny caches per lanemanage on-chip reuse explicitly instead of trusting a cache
lives behind a slow host linkkeep data resident on the device instead of shuttling it back

Each of those has its own page later in this section. They are all the same bet, seen from five angles.

the workload the bet was made for

A neural network's forward pass is, almost entirely, a stack of matrix multiplications. Every output element is an independent dot product of a row and a column — thousands to millions of multiply-accumulates with no ordering constraints between them, all needing the same instruction applied to different data. That is precisely the shape of work the table above asks for, and it is not really a coincidence that deep learning became practical when it did: the hardware that made it possible had been built for rasterizing triangles and was already sitting in every gaming PC.

Training is exactly this, and so is the prefill phase of inference — the pass over the prompt you typed, before any token comes back. Both are large batched matmuls with enough independent work to keep every lane busy, and a GPU runs them at a serious fraction of its peak. This is the bet paying off as designed.

Generation is a different story, and it is worth being precise here because the popular version has it backwards. An LLM emitting text one token at a time is not compute-bound at all. To produce a single token it must read every weight in the model out of memory, do comparatively little arithmetic with each one, and discard it. What limits you is bandwidth, not FLOPs — which is why memory bandwidth predicts LLM performance better than any TFLOPs figure, and why a datacenter GPU decoding one token at a time is running at a small fraction of its arithmetic ceiling. Large language models are not fast because of the GPU. They are tolerable because of it, and most of the engineering in modern inference — caching past keys and values, batching many requests so each weight read serves more than one of them, fusing attention kernels, quantizing weights to move fewer bytes — exists to close the gap between the two stories in this section.

when the GPU is the wrong tool

The honest flip side, since most introductions skip it. A GPU is a poor choice when the work is sequential and each step depends on the previous one; when the problem is too small for the launch overhead to amortize; when control flow is heavily data-dependent and every lane wants a different branch; when the working set is pointer-chasing and irregular rather than arrays; or when the data would spend more time crossing the host link than it would spend being computed on. A CPU with a good cache hierarchy is genuinely the better machine for all of those, and reaching for a GPU anyway is one of the more common ways to spend a week and lose performance.

see the bet on your own machine

Worth doing once, side by side — the contrast in the output is the whole page in concrete numbers.

lscpu | https://man7.org/linux/man-pages/man1/lscpu.1.html | core count, thread count and cache sizes for the host CPU — note how much cache there is per core |'cvg_lscpu'
nvidia-smi --query-gpu=name,memory.total,clocks.max.sm --format=csv | https://developer.nvidia.com/nvidia-system-management-interface | the Nvidia side: memory capacity and peak SM clock, both lower than you might expect next to a CPU's clock |'cvg_nvsmi'
rocm-smi --showproductname --showmeminfo vram | https://rocm.docs.amd.com/projects/rocm_smi_lib/en/latest/ | the same idea on AMD hardware |'cvg_rocmsmi'

related topics

Anatomy of a GPU — what the throughput bet looks like as physical hardware: die, package and memory.
CUDA & HIP — the programming model you use to actually express thousands of independent work items.
GPU Optimization — what to do once a kernel runs but runs badly.
GPU Memory Planning for LLMs — the bandwidth and capacity story above, turned into a budget you can actually compute.

reference

NVIDIA CUDA C++ Programming Guide
AMD HIP documentation
NVIDIA GPU architecture whitepapers
AMD CDNA architecture