SM vs CU: The Core Compute Block
Nvidia stamps out SMs, AMD stamps out CUs. Same idea, different arithmetic.
A GPU die is one compute block repeated many times. Nvidia calls that block a Streaming Multiprocessor, AMD calls it a Compute Unit, and they are the same idea: a self-contained bundle of arithmetic lanes, schedulers, registers and on-chip scratchpad that can run many threads at once. Your kernel's thread blocks get handed out to these, one block per unit, and a GPU scales by having more of them.
They are not, however, directly comparable units, and the marketing numbers built on top of them are comparable even less.
Strip away the vendor names and both contain the same categories of thing:
| Part | What it does |
|---|---|
| Arithmetic lanes | The units that actually do the FP32, FP64 and integer work, grouped into fixed-width SIMD hardware. |
| Schedulers | Pick a ready group of threads each cycle and issue one instruction for the whole group. |
| Register file | Private per-thread storage, and by far the largest on-chip memory. Tens to hundreds of KB per unit. |
| Scratchpad | Programmer-managed shared memory (Nvidia) or LDS (AMD), used for cooperation inside a block. |
| L1 / texture cache | Automatic caching of global memory traffic, sometimes sharing physical storage with the scratchpad. |
| Matrix units | Tensor Cores or Matrix Cores, a separate datapath for small matrix multiply-accumulate. |
| Load/store and special-function units | Address generation, and transcendentals like exp and rsqrt. |
The interesting differences are in the proportions, not the parts list.
| Nvidia SM (Hopper class) | AMD CU (CDNA class) | |
|---|---|---|
| Execution group | Warp of 32 threads | Wavefront of 64 threads |
| Internal structure | Four processing blocks, each with its own scheduler and 32 FP32 lanes | Four SIMD units, each 16 lanes wide, issuing a wave64 over four cycles |
| FP32 lanes per unit | 128 | 64 |
| Scratchpad | Up to 256 KB combined L1 and shared memory, split configurably | 64 KB LDS, separate from L1 |
| Matrix hardware | Tensor Cores | Matrix Cores, driven by MFMA instructions |
| Typical count per die | Roughly 100 to 150 | Roughly 200 to 300 |
Read the last row together with the row above it. AMD ships more units with fewer lanes each, Nvidia ships fewer units with more lanes each, and the totals land closer together than either number suggests on its own.
Both vendors publish a headline core count. Nvidia multiplies SMs by FP32 lanes and calls the result CUDA cores. AMD multiplies CUs by 64 and calls the result stream processors. Neither number is wrong, and comparing them across vendors is still meaningless, for three reasons.
First, a "core" here is a lane, not anything resembling a CPU core. It has no independent program counter and cannot run a different instruction from its neighbours. Second, the clock differs, so equal lane counts do not imply equal throughput. Third, and most importantly, the figure ignores the matrix units entirely, and on any modern AI workload those supply the overwhelming majority of the arithmetic.
The numbers that survive the comparison are the boring ones: achieved FLOPs at the precision you actually use, and memory bandwidth. Prefer them.
Take the two datacenter parts this site uses most, an Nvidia H100 SXM and an AMD MI300X, and follow the headline numbers down from the unit count.
| H100 SXM | MI300X | MI300X ÷ H100 | |
|---|---|---|---|
| SMs / CUs | 132 | 304 | 2.3× |
| FP32 lanes per unit | 128 | 64 | 0.5× |
| Headline "cores" | 16,896 CUDA cores | 19,456 stream processors | 1.15× |
| Peak clock | 1.98 GHz | 2.1 GHz | 1.06× |
| FP32 vector peak | 67 TFLOPS | 163 TFLOPS | 2.4× |
| BF16 matrix peak, dense | 989 TFLOPS | 1,307 TFLOPS | 1.3× |
The core counts differ by 15%. The FP32 figures differ by a factor of 2.4, and the matrix figures by 1.3. The core count predicts neither.
The H100 row is the plain formula from FLOPs and TFLOPS: 16,896 lanes × 2 FLOPs per fused multiply-add × 1.98 GHz ≈ 67 TFLOPS. The same formula gives the MI300X only 82 TFLOPS. The other factor of two comes from packed FP32 instructions, which make each lane do two FP32 operations at once. That doubling is real, but only for code the compiler can pack, which is why a spec sheet is a ceiling and not a forecast. The matrix row depends on neither lanes nor cores; it is set by the Tensor Cores and Matrix Cores, which the "core" count leaves out entirely.
Because a thread block is assigned to exactly one unit and stays there, everything that block needs has to be carved out of that unit's fixed supply. Registers, scratchpad and thread slots are all finite per unit, and whichever runs out first caps how many blocks can be resident at once.
That cap is what occupancy measures, and occupancy is how the machine hides memory latency: the more resident groups, the more likely one is ready to issue while others wait. This is the point where a design decision inside your kernel, say using a few more registers per thread, quietly reduces the number of blocks per unit and costs you performance somewhere else entirely.
It is also why a kernel that launches ten blocks on a two-hundred-unit GPU is slow no matter how good the kernel is. Most of the machine is simply not participating.
The companion sample prints the per-unit budget of your own GPU, then asks the runtime how many 256-thread blocks of one kernel fit on a single unit as the block's scratchpad request grows. The core of it is one call:
int blocksPerUnit = 0;
hipOccupancyMaxActiveBlocksPerMultiprocessor(&blocksPerUnit, scratchpadKernel,
256, scratchpadBytes);
On a small Nvidia card, built as CUDA, it reported 96 KB of scratchpad per SM and this:
| Scratchpad per block | Blocks per SM | What ran out |
|---|---|---|
| 0 to 8 KB | 8 | Thread slots: 8 × 256 = 2,048, the SM's limit |
| 16 KB | 6 | Scratchpad: 6 × 16 = 96 KB |
| 24 KB | 4 | Scratchpad: 4 × 24 = 96 KB |
| 32 KB | 3 | Scratchpad: 3 × 32 = 96 KB |
| 48 KB | 2 | Scratchpad: 2 × 48 = 96 KB |
The first rows cost nothing, because a different resource was already the limit. Past 8 KB, every extra kilobyte of scratchpad costs resident blocks, and the drop comes in steps rather than smoothly. On an MI300X the budget is different, 64 KB of LDS per CU, but the shape of the table is the same.
hipcc and notes for nvcc:
TopNotchNote/gpu/sm_vs_cu_resource_budget.hip
Both vendors report the unit count and its resource limits directly.
multiProcessorCount) and size the grid to several blocks per unit.warpSize.