Warps vs Wavefronts: How GPUs Actually Execute
Threads do not run individually. They run in fixed-size groups, and the size is 32 or 64.
A GPU does not schedule threads one at a time. It schedules them in fixed-size groups that share a single instruction stream: 32 threads on Nvidia, called a warp, and 64 on AMD's CDNA parts, called a wavefront. Every instruction the scheduler issues is issued for the whole group. The group is the real unit of execution on the machine, and the thread you write code for is a lane inside it.
Almost everything that looks strange about GPU performance follows from that, including why an if statement can cost you double and why 32 versus 64 quietly rewrites your tile arithmetic.
Fetching and decoding an instruction costs real silicon and real power. A CPU pays that cost per instruction stream and compensates with caches and prediction. A GPU amortizes it instead: fetch once, decode once, then apply the result to 32 or 64 sets of operands. The control hardware is paid for once and the arithmetic hardware is what scales.
That is the entire economy of the design. It buys enormous arithmetic density, and the bill comes due whenever the lanes in a group want to do different things.
| Hardware | Group | Size |
|---|---|---|
| Nvidia, every architecture to date | Warp | 32 |
| AMD CDNA (MI series, datacenter) | Wavefront | 64 |
| AMD RDNA (Radeon, consumer) | Wavefront | 32 or 64, selectable per kernel |
| Intel Xe | Sub-group | 8, 16 or 32 depending on the kernel |
For anyone writing with AMD datacenter hardware as the target, 64 is the number that matters, and it is not a detail you can paper over at the end. It appears in your block sizes, in how much scratchpad a group needs, in how many groups fit per CU, and in every tile dimension you pick for a matrix kernel. Porting a carefully tuned 32-lane kernel to a 64-lane machine means redoing that arithmetic, not recompiling.
The classic description is that a group shares one program counter. That was literally true for a long time and is worth keeping as the mental model, but modern Nvidia hardware since Volta gives each thread its own program counter, which lets diverged lanes make independent forward progress and avoids some old deadlock patterns.
What did not change is the issue model: the scheduler still issues one instruction for one group at a time, so lanes that are not on the current path are simply masked off and do no useful work. The practical consequence is unchanged too. Code that assumed lanes within a group stay implicitly in step is no longer safe, which is why explicit warp-level synchronization exists and why the old implicitly-synchronized reduction tricks were deprecated.
When lanes in one group take different branches, the hardware does not split the group. It runs one path with the non-participating lanes disabled, then the other path with the mask inverted.
So the cost of a divergent branch is the sum of the paths taken by any lane in the group, not the longest one and certainly not the average. A chain of mutually exclusive branches that each catch a few lanes can serialize into something far slower than the code looks.
The important qualifier is that divergence only matters within a group. Two different groups taking different branches cost nothing extra, because they were always going to be issued separately. This is why the usual advice is not "avoid branches" but "arrange your data so that lanes in the same group agree", which is a data layout problem rather than a control flow one.
A block is carved into whole groups. Ask for 100 threads per block on Nvidia and you get four warps, the last of which runs with 28 of its 32 lanes permanently idle. You paid for 128 thread slots and wasted 28 of them for the lifetime of every block.
The rule that falls out is simple: make block sizes a multiple of the group size. 128, 256 or 512 are safe on both vendors, since both 32 and 64 divide them. The trap is a block of 32, which is a full warp on Nvidia and half a wavefront on CDNA, so a kernel tuned on one and moved to the other silently halves its efficiency in that dimension.
The group size is reported by the runtime, and the profilers will tell you what fraction of your lanes were actually doing work.