the short version

A GPU die is one compute block repeated many times. Nvidia calls that block a Streaming Multiprocessor, AMD calls it a Compute Unit, and they are the same idea: a self-contained bundle of arithmetic lanes, schedulers, registers and on-chip scratchpad that can run many threads at once. Your kernel's thread blocks get handed out to these, one block per unit, and a GPU scales by having more of them.

They are not, however, directly comparable units, and the marketing numbers built on top of them are comparable even less.

what is inside one

Strip away the vendor names and both contain the same categories of thing:

PartWhat it does
Arithmetic lanesThe units that actually do the FP32, FP64 and integer work, grouped into fixed-width SIMD hardware.
SchedulersPick a ready group of threads each cycle and issue one instruction for the whole group.
Register filePrivate per-thread storage, and by far the largest on-chip memory. Tens to hundreds of KB per unit.
ScratchpadProgrammer-managed shared memory (Nvidia) or LDS (AMD), used for cooperation inside a block.
L1 / texture cacheAutomatic caching of global memory traffic, sometimes sharing physical storage with the scratchpad.
Matrix unitsTensor Cores or Matrix Cores, a separate datapath for small matrix multiply-accumulate.
Load/store and special-function unitsAddress generation, and transcendentals like exp and rsqrt.

The interesting differences are in the proportions, not the parts list.

the two, side by side

Nvidia SM (Hopper class)AMD CU (CDNA class)
Execution groupWarp of 32 threadsWavefront of 64 threads
Internal structureFour processing blocks, each with its own scheduler and 32 FP32 lanesFour SIMD units, each 16 lanes wide, issuing a wave64 over four cycles
FP32 lanes per unit12864
ScratchpadUp to 256 KB combined L1 and shared memory, split configurably64 KB LDS, separate from L1
Matrix hardwareTensor CoresMatrix Cores, driven by MFMA instructions
Typical count per dieRoughly 100 to 150Roughly 200 to 300

Read the last row together with the row above it. AMD ships more units with fewer lanes each, Nvidia ships fewer units with more lanes each, and the totals land closer together than either number suggests on its own.

One Hopper SM next to one CDNA 3 CU Both units are split into four quarters. Each quarter of the SM has a scheduler, 32 FP32 lanes, one Tensor Core and 64 KB of registers, and the SM shares up to 256 KB of combined L1 and shared memory. Each quarter of the CU is a SIMD unit with 16 lanes that runs a 64-thread wavefront over four cycles, one Matrix Core and 128 KB of vector registers, and the CU has a separate 64 KB LDS and 32 KB L1. Nvidia SM (Hopper) scheduler 32 FP32 lanes 1 Tensor Core 64 KB registers scheduler 32 FP32 lanes 1 Tensor Core 64 KB registers scheduler 32 FP32 lanes 1 Tensor Core 64 KB registers scheduler 32 FP32 lanes 1 Tensor Core 64 KB registers up to 256 KB L1 + shared memory AMD CU (CDNA 3) SIMD unit 16 lanes × 4 cycles 1 Matrix Core 128 KB registers SIMD unit 16 lanes × 4 cycles 1 Matrix Core 128 KB registers SIMD unit 16 lanes × 4 cycles 1 Matrix Core 128 KB registers SIMD unit 16 lanes × 4 cycles 1 Matrix Core 128 KB registers 64 KB LDS 32 KB L1
Same four-way split, different proportions. The SM has twice the FP32 lanes; the CU has twice the registers per unit and keeps its scratchpad separate from L1. A GPU is simply one of these repeated: 132 times on an H100 SXM, 304 times on an MI300X.

why core counts do not compare

Both vendors publish a headline core count. Nvidia multiplies SMs by FP32 lanes and calls the result CUDA cores. AMD multiplies CUs by 64 and calls the result stream processors. Neither number is wrong, and comparing them across vendors is still meaningless, for three reasons.

First, a "core" here is a lane, not anything resembling a CPU core. It has no independent program counter and cannot run a different instruction from its neighbours. Second, the clock differs, so equal lane counts do not imply equal throughput. Third, and most importantly, the figure ignores the matrix units entirely, and on any modern AI workload those supply the overwhelming majority of the arithmetic.

The numbers that survive the comparison are the boring ones: achieved FLOPs at the precision you actually use, and memory bandwidth. Prefer them.

two real chips, worked through

Take the two datacenter parts this site uses most, an Nvidia H100 SXM and an AMD MI300X, and follow the headline numbers down from the unit count.

H100 SXMMI300XMI300X ÷ H100
SMs / CUs1323042.3×
FP32 lanes per unit128640.5×
Headline "cores"16,896 CUDA cores19,456 stream processors1.15×
Peak clock1.98 GHz2.1 GHz1.06×
FP32 vector peak67 TFLOPS163 TFLOPS2.4×
BF16 matrix peak, dense989 TFLOPS1,307 TFLOPS1.3×

The core counts differ by 15%. The FP32 figures differ by a factor of 2.4, and the matrix figures by 1.3. The core count predicts neither.

The H100 row is the plain formula from FLOPs and TFLOPS: 16,896 lanes × 2 FLOPs per fused multiply-add × 1.98 GHz ≈ 67 TFLOPS. The same formula gives the MI300X only 82 TFLOPS. The other factor of two comes from packed FP32 instructions, which make each lane do two FP32 operations at once. That doubling is real, but only for code the compiler can pack, which is why a spec sheet is a ceiling and not a forecast. The matrix row depends on neither lanes nor cores; it is set by the Tensor Cores and Matrix Cores, which the "core" count leaves out entirely.

the unit is a resource budget

Because a thread block is assigned to exactly one unit and stays there, everything that block needs has to be carved out of that unit's fixed supply. Registers, scratchpad and thread slots are all finite per unit, and whichever runs out first caps how many blocks can be resident at once.

That cap is what occupancy measures, and occupancy is how the machine hides memory latency: the more resident groups, the more likely one is ready to issue while others wait. This is the point where a design decision inside your kernel, say using a few more registers per thread, quietly reduces the number of blocks per unit and costs you performance somewhere else entirely.

It is also why a kernel that launches ten blocks on a two-hundred-unit GPU is slow no matter how good the kernel is. Most of the machine is simply not participating.

try it

The companion sample prints the per-unit budget of your own GPU, then asks the runtime how many 256-thread blocks of one kernel fit on a single unit as the block's scratchpad request grows. The core of it is one call:


int blocksPerUnit = 0;
hipOccupancyMaxActiveBlocksPerMultiprocessor(&blocksPerUnit, scratchpadKernel,
                                             256, scratchpadBytes);
            

On a small Nvidia card, built as CUDA, it reported 96 KB of scratchpad per SM and this:

Scratchpad per blockBlocks per SMWhat ran out
0 to 8 KB8Thread slots: 8 × 256 = 2,048, the SM's limit
16 KB6Scratchpad: 6 × 16 = 96 KB
24 KB4Scratchpad: 4 × 24 = 96 KB
32 KB3Scratchpad: 3 × 32 = 96 KB
48 KB2Scratchpad: 2 × 48 = 96 KB

The first rows cost nothing, because a different resource was already the limit. Past 8 KB, every extra kilobyte of scratchpad costs resident blocks, and the drop comes in steps rather than smoothly. On an MI300X the budget is different, 64 KB of LDS per CU, but the shape of the table is the same.

The full program, with build instructions for hipcc and notes for nvcc: TopNotchNote/gpu/sm_vs_cu_resource_budget.hip

see it on your own machine

Both vendors report the unit count and its resource limits directly.

nvidia-smi --query-gpu=name,count --format=csv | https://developer.nvidia.com/nvidia-system-management-interface | confirms which device you are querying before you go looking for its SM count |'smcu_smi'
rocminfo | https://rocm.docs.amd.com/projects/rocminfo/en/latest/ | lists each AMD agent with its compute unit count, wavefront size and LDS size per CU |'smcu_rocminfo'
hipcc --offload-arch=native -O3 kernel.hip -o kernel | https://rocm.docs.amd.com/projects/HIP/en/latest/ | builds for the gfx target actually installed, which determines the CU layout you are compiling against |'smcu_hipcc'

rules of thumb

  • Never compare GPUs by core count. Compare achieved FLOPs at the precision you use, and memory bandwidth.
  • Read the SM/CU count at runtime (multiProcessorCount) and size the grid to several blocks per unit.
  • Know your per-unit budget: registers, scratchpad and thread slots. The sample prints all three.
  • Before adding scratchpad or registers to a block, check what it does to blocks per unit with the occupancy API or a profiler.
  • Do not hardcode 32 or 64 as the group size; read warpSize.

related topics

Warps vs Wavefronts — the execution group these units schedule, and why 32 versus 64 changes your tile sizes.
The GPU Memory Hierarchy — where the register file and scratchpad sit relative to everything else.
What a Kernel Launch Physically Does — how blocks reach these units in the first place.
Occupancy and Register Pressure — the resource budget above, as the number to tune.
Nvidia ↔ AMD GPU Glossary — every other name that differs between the two.

reference

NVIDIA GPU architecture whitepapers
AMD CDNA architecture
AMD HIP documentation
NVIDIA CUDA C++ Programming Guide