CPU vs GPU: Why GPUs Exist
One design bet — throughput over latency — explains almost everything else.
A CPU is built to finish one task as soon as possible. A GPU is built to finish a million tasks per second, and does not care how long any single one of them takes. Both chips get roughly the same budget — a few hundred square millimetres of silicon and a power envelope — and they spend it in opposite directions. Everything else on this page, and most of what makes GPU code fast or slow, follows from that one choice.
The interesting comparison is not "how many cores" but what the silicon is spent on. A CPU spends most of its area on machinery that makes a single instruction stream run fast: large caches, out-of-order scheduling, register renaming, branch predictors, deep pipelines. None of that machinery does arithmetic — it exists to keep the few arithmetic units fed. A GPU deletes almost all of it and spends the area on arithmetic units instead.
| Budget spent on | CPU | GPU |
|---|---|---|
| Independent cores | Tens (8–128 on servers) | Dozens to hundreds of SMs/CUs, each with many lanes |
| Arithmetic lanes | Hundreds to a couple thousand, via SIMD (AVX-512, SVE) | Tens of thousands |
| Cache per lane | Very large (tens of MB of L3) | Very small (tens to a couple hundred KB on-chip per SM/CU) |
| Out-of-order execution | Yes, aggressive | No — in-order issue per warp/wavefront |
| Branch prediction | Yes, sophisticated | Essentially none |
| Clock speed | High (4–6 GHz boost) | Lower (roughly 1–2.5 GHz) |
| Memory priority | Low latency | High bandwidth |
Notice the last row especially. A CPU's memory system is tuned to return one cache line as quickly as possible; a GPU's is tuned to move enormous volumes of data per second even though any individual access is slower than the CPU's. Those are not the same goal, and you cannot maximize both.
Here is the part that is easy to miss. A GPU does not make memory faster — it makes memory slower. A read that misses cache and goes out to GPU memory takes on the order of 400–600 nanoseconds, against roughly 80–100 for a CPU reaching its own DRAM. Like for like, at the same level of the hierarchy, the GPU loses. It wins anyway because of what it does while waiting.
A CPU facing a slow memory access tries to avoid the stall: it predicts the branch, prefetches the line, reorders later instructions to run ahead. All of that costs transistors and power. A GPU facing the same stall simply switches to another group of threads that is ready to run, and issues their instruction this cycle instead. Each SM/CU keeps far more threads resident than it can execute at once, precisely so that there is always someone else to run.
This is why "just add more parallelism" is the standard GPU fix for a slow kernel, and why a GPU running only a handful of threads is catastrophically slow — not because the threads are slow, but because there is nobody left to hide behind. It is also the reason occupancy becomes a thing you tune later on.
Once you accept the throughput bet, most GPU performance advice stops being a list to memorize and becomes derivable:
| Because the GPU… | …your code must |
|---|---|
| hides latency with spare threads | expose thousands of independent work items, not dozens |
| executes lanes in lockstep groups | avoid divergent branches inside a group |
| optimizes for bandwidth, not latency | read memory in wide, contiguous, neighbour-friendly patterns |
| has tiny caches per lane | manage on-chip reuse explicitly instead of trusting a cache |
| lives behind a slow host link | keep data resident on the device instead of shuttling it back |
Each of those has its own page later in this section. They are all the same bet, seen from five angles.
A neural network's forward pass is, almost entirely, a stack of matrix multiplications. Every output element is an independent dot product of a row and a column — thousands to millions of multiply-accumulates with no ordering constraints between them, all needing the same instruction applied to different data. That is precisely the shape of work the table above asks for, and it is not really a coincidence that deep learning became practical when it did: the hardware that made it possible had been built for rasterizing triangles and was already sitting in every gaming PC.
Training is exactly this, and so is the prefill phase of inference — the pass over the prompt you typed, before any token comes back. Both are large batched matmuls with enough independent work to keep every lane busy, and a GPU runs them at a serious fraction of its peak. This is the bet paying off as designed.
Generation is a different story, and it is worth being precise here because the popular version has it backwards. An LLM emitting text one token at a time is not compute-bound at all. To produce a single token it must read every weight in the model out of memory, do comparatively little arithmetic with each one, and discard it. What limits you is bandwidth, not FLOPs — which is why memory bandwidth predicts LLM performance better than any TFLOPs figure, and why a datacenter GPU decoding one token at a time is running at a small fraction of its arithmetic ceiling. Large language models are not fast because of the GPU. They are tolerable because of it, and most of the engineering in modern inference — caching past keys and values, batching many requests so each weight read serves more than one of them, fusing attention kernels, quantizing weights to move fewer bytes — exists to close the gap between the two stories in this section.
The honest flip side, since most introductions skip it. A GPU is a poor choice when the work is sequential and each step depends on the previous one; when the problem is too small for the launch overhead to amortize; when control flow is heavily data-dependent and every lane wants a different branch; when the working set is pointer-chasing and irregular rather than arrays; or when the data would spend more time crossing the host link than it would spend being computed on. A CPU with a good cache hierarchy is genuinely the better machine for all of those, and reaching for a GPU anyway is one of the more common ways to spend a week and lose performance.
Worth doing once, side by side — the contrast in the output is the whole page in concrete numbers.