Nvidia ↔ AMD GPU Glossary
Two vendors, one set of ideas, two sets of names. The translation table.
Nvidia and AMD GPUs are built on the same ideas: a large number of identical compute blocks, threads grouped into fixed-size bundles that execute in lockstep, a small fast scratchpad per block, and a wide memory system behind it all. Almost every concept has a name on each side, and the names rarely match. Most confusion when switching between CUDA and HIP, or reading the other vendor's documentation, is vocabulary rather than substance.
This page is the lookup table. Each row links to the page on this site that explains the concept, where one exists. Where the two sides genuinely differ, the row says so rather than pretending they are the same thing.
| Nvidia / CUDA | AMD / HIP | What it is |
|---|---|---|
| Thread | Work-item (HIP code still says thread) | One instance of the kernel, with its own registers and index. |
| Warp, 32 threads | Wavefront, 64 threads on CDNA datacenter parts; 32 or 64 on RDNA consumer parts | The group that executes one instruction at a time in lockstep. See Warps vs Wavefronts. |
| Thread block, or CTA | Workgroup (HIP code still says block) | Threads that land on one compute block together and can share its scratchpad and barrier. See What a Kernel Launch Physically Does. |
| Grid | Grid (NDRange in OpenCL) | All the blocks of one launch. |
| Thread block cluster | No direct equivalent | Hopper and later: a group of blocks guaranteed to run at once on neighbouring SMs and able to read each other's shared memory. |
| Stream | Stream | An ordered queue of GPU work. See Streams. |
| CUDA Graph | HIP Graph | A recorded sequence of launches replayed with one submission. |
| Nvidia | AMD | What it is |
|---|---|---|
| SM, Streaming Multiprocessor | CU, Compute Unit | The repeated compute block that thread blocks are assigned to. See SM vs CU. |
| SM sub-partition (4 per SM) | SIMD unit (4 per CU) | A quarter of the block with its own scheduler and register file; each warp or wavefront lives on one. |
| CUDA core | Stream processor | A single FP32 lane. A marketing count, not a core in the CPU sense. |
| Tensor Core | Matrix Core (MFMA instructions) | Units that do a small matrix multiply per instruction. See Tensor Cores vs Matrix Cores. |
| GPC, Graphics Processing Cluster | Shader Engine; on MI300, the XCD | A group of SMs/CUs sharing front-end hardware. An MI300X has 8 XCDs, separate dies of 38 active CUs each. |
| NVLink | Infinity Fabric (xGMI) | The direct GPU-to-GPU link, much faster than PCIe. |
| NVSwitch | No switch in the 8-GPU MI300X platform | Nvidia routes GPU-to-GPU traffic through switch chips; the MI300X platform connects every GPU directly to every other. |
Compute capability, e.g. sm_90 | GFX target, e.g. gfx942 | The hardware generation code you compile for. See Compiling GPU Code. |
| Nvidia / CUDA | AMD / HIP | What it is |
|---|---|---|
| Registers | VGPRs and SGPRs | Per-thread storage. AMD splits it into vector registers (one value per lane) and scalar registers (one value shared by the whole wavefront). |
| Shared memory | LDS, Local Data Share | The fast on-chip scratchpad shared by one block. See Shared Memory & Bank Conflicts. |
| Local memory | Scratch (private) memory | Per-thread spill space when registers run out. Despite the name, it lives in device memory and is slow. |
| Global memory | Global memory | Device memory, HBM or GDDR, visible to every thread. See The GPU Memory Hierarchy. |
| L2 cache | L2 cache, plus Infinity Cache on some parts | The last on-chip level before device memory. MI300X adds a 256 MB Infinity Cache below its L2s. |
Unified memory, cudaMallocManaged | hipMallocManaged | One pointer valid on host and device, migrated on demand. |
| Nvidia | AMD | What it is |
|---|---|---|
| CUDA | HIP (the language and API), ROCm (the whole platform) | The programming model. See CUDA & HIP. |
nvcc | hipcc, amdclang++ | The compiler driver. |
| PTX | No equivalent | Nvidia's portable virtual ISA, compiled to machine code at build or load time. AMD compiles straight to the machine ISA of each GFX target. |
| SASS | AMDGPU ISA (CDNA / RDNA assembly) | The real machine instructions. See Assembly & Low-Level. |
| Driver and CUDA runtime | amdgpu kernel driver, ROCr, HIP runtime | The layers under your program. See Driver, Toolkit and Runtime. |
| cuBLAS, cuBLASLt | rocBLAS, hipBLASLt | Dense linear algebra, above all matrix multiply. See GPU Libraries. |
| cuDNN | MIOpen | Deep learning primitives: convolution, normalization, attention. |
| NCCL | RCCL | Collective communication across GPUs: all-reduce, all-gather. |
| CUTLASS | Composable Kernel (CK) | Template libraries for building your own high-performance kernels. |
| Thrust, CUB | rocThrust, hipCUB, rocPRIM | Parallel algorithms: sort, scan, reduce. |
| cuFFT, cuSPARSE, cuRAND | rocFFT, rocSPARSE, rocRAND (plus hip* wrappers) | FFTs, sparse linear algebra, random numbers. |
AMD libraries come in pairs more often than not. The roc* library is the AMD implementation; the hip* library is a thin portable interface with the same API shape as the CUDA library, which calls the roc* one on AMD hardware and the CUDA one on Nvidia. Code ported from CUDA usually targets the hip* name. See Porting CUDA to HIP.
| Nvidia | AMD | What it is for |
|---|---|---|
nvidia-smi | amd-smi, rocm-smi | Device status: clocks, power, memory, utilization. |
deviceQuery sample | rocminfo | Every property of every device. |
| Nsight Systems | rocprofv3, ROCm Systems Profiler | Whole-program timelines. See GPU Profiling Tools. |
| Nsight Compute | ROCm Compute Profiler | Deep per-kernel analysis. |
cuda-gdb | rocgdb | A debugger that can stop inside kernels. |
A few names exist on both sides and mean different things. These cause more wrong conclusions than the ones that simply differ.
Local memory. In CUDA it is the slow, per-thread spill area in device memory. In OpenCL, and in much AMD documentation, "local memory" means the fast on-chip scratchpad, the LDS, which CUDA calls shared memory. Read "local memory is slow" in the wrong vendor's context and you conclude the opposite of the truth.
Core. A "CUDA core" and a "stream processor" are each one FP32 lane, and the two vendors count them differently, so core counts cannot be compared across vendors. Neither is a core in the CPU sense; the SM or CU is the closer analogue. SM vs CU goes through the arithmetic.
Wavefront size. Code that assumes 64 threads per group is correct on AMD's datacenter GPUs and wrong on its consumer ones, which default to 32. Code that assumes 32 is wrong on the datacenter parts. Query warpSize at runtime instead of hardcoding either.
Shader Engine and SM. Both sound like the basic unit. The SM is Nvidia's compute block; a Shader Engine is a much larger AMD grouping of many CUs, closer to Nvidia's GPC.