CUDA & HIP
What the two toolchains actually are, how they relate, and where the portability boundary really sits.
CUDA is Nvidia's platform for general-purpose GPU programming: a set of C++ language extensions, a compiler, a runtime, and a stack of libraries. HIP is AMD's equivalent, and it was deliberately designed to look almost exactly like CUDA so that code can move between them.
That similarity is the whole strategy. HIP is not an independent design that happens to be comparable; it mirrors CUDA's vocabulary on purpose, so that a CUDA programmer already knows HIP and an existing codebase can be translated largely mechanically.
The word gets used for at least four things, which is worth untangling early.
| The word usually means | Which is |
|---|---|
| The language extensions | __global__, __device__, __shared__ and the launch syntax, on top of C++ |
| The toolkit | nvcc, headers, and the development libraries you build against |
| The runtime | The library your binary calls at execution time to allocate, copy and launch |
| The library ecosystem | cuBLAS, cuDNN, cuFFT, NCCL and the rest, built on top of the above |
Those have separate version numbers and are allowed to disagree, which causes enough confusion to deserve its own page.
ROCm is AMD's overall software stack, the counterpart to the CUDA toolkit plus its libraries. HIP is the programming interface within it: the language extensions and runtime API that mirror CUDA's.
The mirroring is close enough to be mechanical. cudaMalloc becomes hipMalloc, cudaMemcpy becomes hipMemcpy, cublasSgemm becomes hipblasSgemm. Inside a kernel, threadIdx, blockIdx, blockDim, __syncthreads and __shared__ are spelled identically in both.
One consequence that surprises people: HIP compiles for Nvidia hardware too, using nvcc underneath. So HIP is the portable choice rather than the AMD-only one, which is the main argument for writing new portable code in it.
It is not where you would guess. The kernel body ports almost for free, because the language extensions are the same. Nearly all of the difference is in host-side API calls, and those are a rename.
The genuinely hard part is neither of those. It is the execution group size: 32 threads on Nvidia, 64 on AMD CDNA. Any kernel that assumed 32, whether through a hardcoded literal, a shuffle-based reduction, or a 32-bit ballot mask, compiles cleanly on AMD and produces wrong answers. No translation tool catches it, because nothing in that code mentions a CUDA function.
So the honest summary is that the API ports mechanically and the assumptions do not. Porting CUDA to HIP goes through exactly what to grep for.
| Situation | Reasonable choice |
|---|---|
| Nvidia hardware only, and you want the newest features first | CUDA. New capabilities land there first and the ecosystem is deeper. |
| AMD hardware, now or later | HIP. It is the native interface and there is no reason to add a translation step. |
| You want one source tree for both | HIP, which compiles for both vendors. |
| You are writing kernels through a framework | Often neither directly. PyTorch, Triton and the vendor libraries cover a great deal of ground before you need to write a kernel at all. |
Each of the following was a section of this page and is now a page of its own, with room for the code and the failure modes that a summary could not carry.