the short version

Nvidia and AMD GPUs are built on the same ideas: a large number of identical compute blocks, threads grouped into fixed-size bundles that execute in lockstep, a small fast scratchpad per block, and a wide memory system behind it all. Almost every concept has a name on each side, and the names rarely match. Most confusion when switching between CUDA and HIP, or reading the other vendor's documentation, is vocabulary rather than substance.

This page is the lookup table. Each row links to the page on this site that explains the concept, where one exists. Where the two sides genuinely differ, the row says so rather than pretending they are the same thing.

how the work is organized

Nvidia / CUDAAMD / HIPWhat it is
ThreadWork-item (HIP code still says thread)One instance of the kernel, with its own registers and index.
Warp, 32 threadsWavefront, 64 threads on CDNA datacenter parts; 32 or 64 on RDNA consumer partsThe group that executes one instruction at a time in lockstep. See Warps vs Wavefronts.
Thread block, or CTAWorkgroup (HIP code still says block)Threads that land on one compute block together and can share its scratchpad and barrier. See What a Kernel Launch Physically Does.
GridGrid (NDRange in OpenCL)All the blocks of one launch.
Thread block clusterNo direct equivalentHopper and later: a group of blocks guaranteed to run at once on neighbouring SMs and able to read each other's shared memory.
StreamStreamAn ordered queue of GPU work. See Streams.
CUDA GraphHIP GraphA recorded sequence of launches replayed with one submission.

the hardware

NvidiaAMDWhat it is
SM, Streaming MultiprocessorCU, Compute UnitThe repeated compute block that thread blocks are assigned to. See SM vs CU.
SM sub-partition (4 per SM)SIMD unit (4 per CU)A quarter of the block with its own scheduler and register file; each warp or wavefront lives on one.
CUDA coreStream processorA single FP32 lane. A marketing count, not a core in the CPU sense.
Tensor CoreMatrix Core (MFMA instructions)Units that do a small matrix multiply per instruction. See Tensor Cores vs Matrix Cores.
GPC, Graphics Processing ClusterShader Engine; on MI300, the XCDA group of SMs/CUs sharing front-end hardware. An MI300X has 8 XCDs, separate dies of 38 active CUs each.
NVLinkInfinity Fabric (xGMI)The direct GPU-to-GPU link, much faster than PCIe.
NVSwitchNo switch in the 8-GPU MI300X platformNvidia routes GPU-to-GPU traffic through switch chips; the MI300X platform connects every GPU directly to every other.
Compute capability, e.g. sm_90GFX target, e.g. gfx942The hardware generation code you compile for. See Compiling GPU Code.

memory

Nvidia / CUDAAMD / HIPWhat it is
RegistersVGPRs and SGPRsPer-thread storage. AMD splits it into vector registers (one value per lane) and scalar registers (one value shared by the whole wavefront).
Shared memoryLDS, Local Data ShareThe fast on-chip scratchpad shared by one block. See Shared Memory & Bank Conflicts.
Local memoryScratch (private) memoryPer-thread spill space when registers run out. Despite the name, it lives in device memory and is slow.
Global memoryGlobal memoryDevice memory, HBM or GDDR, visible to every thread. See The GPU Memory Hierarchy.
L2 cacheL2 cache, plus Infinity Cache on some partsThe last on-chip level before device memory. MI300X adds a 256 MB Infinity Cache below its L2s.
Unified memory, cudaMallocManagedhipMallocManagedOne pointer valid on host and device, migrated on demand.

software stack and libraries

NvidiaAMDWhat it is
CUDAHIP (the language and API), ROCm (the whole platform)The programming model. See CUDA & HIP.
nvcchipcc, amdclang++The compiler driver.
PTXNo equivalentNvidia's portable virtual ISA, compiled to machine code at build or load time. AMD compiles straight to the machine ISA of each GFX target.
SASSAMDGPU ISA (CDNA / RDNA assembly)The real machine instructions. See Assembly & Low-Level.
Driver and CUDA runtimeamdgpu kernel driver, ROCr, HIP runtimeThe layers under your program. See Driver, Toolkit and Runtime.
cuBLAS, cuBLASLtrocBLAS, hipBLASLtDense linear algebra, above all matrix multiply. See GPU Libraries.
cuDNNMIOpenDeep learning primitives: convolution, normalization, attention.
NCCLRCCLCollective communication across GPUs: all-reduce, all-gather.
CUTLASSComposable Kernel (CK)Template libraries for building your own high-performance kernels.
Thrust, CUBrocThrust, hipCUB, rocPRIMParallel algorithms: sort, scan, reduce.
cuFFT, cuSPARSE, cuRANDrocFFT, rocSPARSE, rocRAND (plus hip* wrappers)FFTs, sparse linear algebra, random numbers.

AMD libraries come in pairs more often than not. The roc* library is the AMD implementation; the hip* library is a thin portable interface with the same API shape as the CUDA library, which calls the roc* one on AMD hardware and the CUDA one on Nvidia. Code ported from CUDA usually targets the hip* name. See Porting CUDA to HIP.

tools

NvidiaAMDWhat it is for
nvidia-smiamd-smi, rocm-smiDevice status: clocks, power, memory, utilization.
deviceQuery samplerocminfoEvery property of every device.
Nsight Systemsrocprofv3, ROCm Systems ProfilerWhole-program timelines. See GPU Profiling Tools.
Nsight ComputeROCm Compute ProfilerDeep per-kernel analysis.
cuda-gdbrocgdbA debugger that can stop inside kernels.

false friends

A few names exist on both sides and mean different things. These cause more wrong conclusions than the ones that simply differ.

Local memory. In CUDA it is the slow, per-thread spill area in device memory. In OpenCL, and in much AMD documentation, "local memory" means the fast on-chip scratchpad, the LDS, which CUDA calls shared memory. Read "local memory is slow" in the wrong vendor's context and you conclude the opposite of the truth.

Core. A "CUDA core" and a "stream processor" are each one FP32 lane, and the two vendors count them differently, so core counts cannot be compared across vendors. Neither is a core in the CPU sense; the SM or CU is the closer analogue. SM vs CU goes through the arithmetic.

Wavefront size. Code that assumes 64 threads per group is correct on AMD's datacenter GPUs and wrong on its consumer ones, which default to 32. Code that assumes 32 is wrong on the datacenter parts. Query warpSize at runtime instead of hardcoding either.

Shader Engine and SM. Both sound like the basic unit. The SM is Nvidia's compute block; a Shader Engine is a much larger AMD grouping of many CUs, closer to Nvidia's GPC.

see it on your own machine

rocminfo | https://rocm.docs.amd.com/projects/rocminfo/en/latest/ | AMD: CU count, wavefront size, GFX target and LDS size for every agent |'gl_rocminfo'
nvidia-smi -q | https://developer.nvidia.com/nvidia-system-management-interface | Nvidia: the full device report, including clocks, memory and compute capability |'gl_smi'
hipcc --version | https://rocm.docs.amd.com/projects/HIP/en/latest/ | which HIP and ROCm version, and which platform, hipcc is compiling for |'gl_hipcc'

related topics

SM vs CU — the most important row in this table, in depth.
Warps vs Wavefronts — the second most important.
Porting CUDA to HIP — putting the software rows to work.

reference

ROCm device hardware glossary
HIPIFY: CUDA to HIP API mapping tables
NVIDIA CUDA C++ Programming Guide
AMD MI300 architecture overview