the short version

The GPU was not designed for you. It was designed to draw triangles, and the parts of it you now use for matrix multiplication are the parts that happened to generalize. Almost everything that feels arbitrary about GPU programming, the lockstep thread groups, the explicit scratchpad, the aversion to branches, is a direct inheritance from a machine built to shade millions of independent pixels.

Knowing the sequence makes the design read as a series of reasonable decisions rather than a pile of quirks.

fixed function, roughly to 2000

Early graphics accelerators had no programmability worth the name. The pipeline was wired in: transform vertices, light them, rasterize into pixels, apply a texture, write out. You configured it through a state machine and it did those steps in that order.

What matters is the shape of the problem it was built around. Every vertex is independent of every other vertex. Every pixel is independent of every other pixel. There are millions of them, nobody cares which finishes first, and the same operation applies to all. A machine optimized for that workload is going to look like throughput hardware with no branch prediction, and it did.

programmable shaders, the early 2000s

Then the fixed stages became programmable. Vertex and pixel shaders let developers supply small programs that ran per vertex or per pixel, initially tiny and heavily restricted, growing quickly in capability. The hardware still executed them across enormous numbers of independent elements, so the execution model did not change: one program, applied in lockstep across many data items.

That is SIMT, and it was invented for shading, not for compute. The warp and the wavefront are what a pixel shader running over a tile of the screen looks like when you write it down in hardware.

the GPGPU hacks, around 2003 to 2006

Researchers noticed that a pixel shader is a general arithmetic engine with an inconvenient interface, and started smuggling real computation through it. You encoded your input as a texture, drew a rectangle so the shader ran once per element, and read your results back as an image. Linear algebra and physics simulation were expressed as rendering operations because rendering was the only way in.

It worked, it was miserable, and it proved the demand. Early research languages tried to hide the graphics layer, and both vendors could see that the interface, not the silicon, was the obstacle.

unification and CUDA, 2006 to 2007

Two things happened close together. The hardware unified: instead of separate vertex and pixel shader units, one pool of general shader cores handled both, which meant there was now a single programmable array with no graphics-specific identity. And the software exposed it directly. CUDA arrived in 2007 and let you write a C-like kernel and launch it over a grid, with no textures and no rectangles.

This is the moment the modern programming model appears, essentially complete. Grids, blocks, threads, shared memory and barriers are all there, and the shape of a kernel you write today would be recognizable to someone from that release. AMD went through its own sequence toward the same destination, arriving at ROCm and HIP, which deliberately mirrors CUDA's vocabulary.

the compute era, and then the AI era

Once the interface existed, the hardware started growing features that graphics never asked for: strong double precision, ECC memory, and a product line sold with no display outputs at all. Scientific computing moved onto GPUs through the late 2000s and 2010s.

Then deep learning arrived and turned out to be, almost entirely, dense matrix multiplication. The hardware responded by adding dedicated matrix units, first shipping in 2017, alongside high-bandwidth memory on the package and fast device-to-device interconnect. Precision went the other way from the scientific era, from FP64 down through FP32, FP16, BF16 and now FP8, because training tolerates narrow inputs and narrow inputs mean more arithmetic per transistor.

The current generation is shaped by one workload in a way the fixed-function era would find familiar. It is again a machine built for a specific problem, and again people are discovering what else it can be made to do.

what the history explains

Feature that seems oddThe historical reason
Threads execute in fixed-size lockstep groupsOne shader program applied across many independent pixels.
Branches are expensiveShading rarely branched, so no transistors were ever spent predicting them.
The scratchpad is managed by handGraphics workloads had predictable access patterns, so a big automatic cache was not worth the area.
No ordering guarantee between blocksPixels were always independent, so nothing was ever promised.
Enormous bandwidth, mediocre latencyTexture fetching is a streaming problem, not a pointer-chasing one.
Precision keeps getting narrowerColour never needed many bits, and neither, it turns out, does a neural network.

related topics

CPU vs GPU: Why GPUs Exist — the design bet this history produced, stated directly.
Warps vs Wavefronts — the shading execution model, still running your kernels.
Tensor Cores vs Matrix Cores — the most recent layer of workload-specific hardware.

reference

NVIDIA CUDA C++ Programming Guide
AMD HIP documentation
NVIDIA GPU architecture whitepapers
Khronos OpenCL