the short version

A GPU does not schedule threads one at a time. It schedules them in fixed-size groups that share a single instruction stream: 32 threads on Nvidia, called a warp, and 64 on AMD's CDNA parts, called a wavefront. Every instruction the scheduler issues is issued for the whole group. The group is the real unit of execution on the machine, and the thread you write code for is a lane inside it.

Almost everything that looks strange about GPU performance follows from that, including why an if statement can cost you double and why 32 versus 64 quietly rewrites your tile arithmetic.

why groups exist at all

Fetching and decoding an instruction costs real silicon and real power. A CPU pays that cost per instruction stream and compensates with caches and prediction. A GPU amortizes it instead: fetch once, decode once, then apply the result to 32 or 64 sets of operands. The control hardware is paid for once and the arithmetic hardware is what scales.

That is the entire economy of the design. It buys enormous arithmetic density, and the bill comes due whenever the lanes in a group want to do different things.

32 or 64, and where each one turns up

HardwareGroupSize
Nvidia, every architecture to dateWarp32
AMD CDNA (MI series, datacenter)Wavefront64
AMD RDNA (Radeon, consumer)Wavefront32 or 64, selectable per kernel
Intel XeSub-group8, 16 or 32 depending on the kernel

For anyone writing with AMD datacenter hardware as the target, 64 is the number that matters, and it is not a detail you can paper over at the end. It appears in your block sizes, in how much scratchpad a group needs, in how many groups fit per CU, and in every tile dimension you pick for a matrix kernel. Porting a carefully tuned 32-lane kernel to a 64-lane machine means redoing that arithmetic, not recompiling.

what lockstep means now

The classic description is that a group shares one program counter. That was literally true for a long time and is worth keeping as the mental model, but modern Nvidia hardware since Volta gives each thread its own program counter, which lets diverged lanes make independent forward progress and avoids some old deadlock patterns.

What did not change is the issue model: the scheduler still issues one instruction for one group at a time, so lanes that are not on the current path are simply masked off and do no useful work. The practical consequence is unchanged too. Code that assumed lanes within a group stay implicitly in step is no longer safe, which is why explicit warp-level synchronization exists and why the old implicitly-synchronized reduction tricks were deprecated.

divergence, and what it costs

When lanes in one group take different branches, the hardware does not split the group. It runs one path with the non-participating lanes disabled, then the other path with the mask inverted.

How a divergent branch executes in one lane group Four stages of an eight-lane group. Before the branch every lane is active. In the if body only the lanes that passed the test are active and the rest are idle. In the else body the pattern inverts. After the branch all lanes are active again, so the branch cost is the sum of both bodies rather than the longer one. before the branch one instruction, every lane working if (cond) { ... } lanes that failed the test sit idle else { ... } now the other half sits idle after the branch converged, back to full width both bodies run in turn
The lanes are never doing two different things at once. The group walks the if body with half its lanes switched off, then the else body with the other half off, so a divergent branch costs both sides added together.

So the cost of a divergent branch is the sum of the paths taken by any lane in the group, not the longest one and certainly not the average. A chain of mutually exclusive branches that each catch a few lanes can serialize into something far slower than the code looks.

The important qualifier is that divergence only matters within a group. Two different groups taking different branches cost nothing extra, because they were always going to be issued separately. This is why the usual advice is not "avoid branches" but "arrange your data so that lanes in the same group agree", which is a data layout problem rather than a control flow one.

the number decides your block size

A block is carved into whole groups. Ask for 100 threads per block on Nvidia and you get four warps, the last of which runs with 28 of its 32 lanes permanently idle. You paid for 128 thread slots and wasted 28 of them for the lifetime of every block.

The rule that falls out is simple: make block sizes a multiple of the group size. 128, 256 or 512 are safe on both vendors, since both 32 and 64 divide them. The trap is a block of 32, which is a full warp on Nvidia and half a wavefront on CDNA, so a kernel tuned on one and moved to the other silently halves its efficiency in that dimension.

see it on your own machine

The group size is reported by the runtime, and the profilers will tell you what fraction of your lanes were actually doing work.

rocminfo | https://rocm.docs.amd.com/projects/rocminfo/en/latest/ | reports wavefront size per agent, alongside compute unit count and LDS size |'ww_rocminfo'
ncu --metrics smsp__thread_inst_executed_per_inst_executed.ratio ./my_app | https://docs.nvidia.com/nsight-compute/ | average active lanes per issued instruction; a value well under 32 means divergence is costing you |'ww_ncu'
rocprofv3 --kernel-trace --stats -- ./my_app | https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/ | per-kernel statistics on the AMD stack, the starting point for the same question |'ww_rocprof'

related topics

SM vs CU — the compute block that schedules these groups, and how many it keeps resident.
The GPU Memory Hierarchy — why a group reading neighbouring addresses is so much faster than one reading scattered ones.
GPU Optimization — divergence and coalescing as things to measure and fix.

reference

CUDA C++ Programming Guide: SIMT architecture
AMD HIP documentation
AMD CDNA architecture
NVIDIA Nsight Compute