the short version

CUDA is Nvidia's platform for general-purpose GPU programming: a set of C++ language extensions, a compiler, a runtime, and a stack of libraries. HIP is AMD's equivalent, and it was deliberately designed to look almost exactly like CUDA so that code can move between them.

That similarity is the whole strategy. HIP is not an independent design that happens to be comparable; it mirrors CUDA's vocabulary on purpose, so that a CUDA programmer already knows HIP and an existing codebase can be translated largely mechanically.

what CUDA is, precisely

The word gets used for at least four things, which is worth untangling early.

The word usually meansWhich is
The language extensions__global__, __device__, __shared__ and the launch syntax, on top of C++
The toolkitnvcc, headers, and the development libraries you build against
The runtimeThe library your binary calls at execution time to allocate, copy and launch
The library ecosystemcuBLAS, cuDNN, cuFFT, NCCL and the rest, built on top of the above

Those have separate version numbers and are allowed to disagree, which causes enough confusion to deserve its own page.

what HIP and ROCm are

ROCm is AMD's overall software stack, the counterpart to the CUDA toolkit plus its libraries. HIP is the programming interface within it: the language extensions and runtime API that mirror CUDA's.

The mirroring is close enough to be mechanical. cudaMalloc becomes hipMalloc, cudaMemcpy becomes hipMemcpy, cublasSgemm becomes hipblasSgemm. Inside a kernel, threadIdx, blockIdx, blockDim, __syncthreads and __shared__ are spelled identically in both.

One consequence that surprises people: HIP compiles for Nvidia hardware too, using nvcc underneath. So HIP is the portable choice rather than the AMD-only one, which is the main argument for writing new portable code in it.

where the portability boundary actually is

It is not where you would guess. The kernel body ports almost for free, because the language extensions are the same. Nearly all of the difference is in host-side API calls, and those are a rename.

The genuinely hard part is neither of those. It is the execution group size: 32 threads on Nvidia, 64 on AMD CDNA. Any kernel that assumed 32, whether through a hardcoded literal, a shuffle-based reduction, or a 32-bit ballot mask, compiles cleanly on AMD and produces wrong answers. No translation tool catches it, because nothing in that code mentions a CUDA function.

So the honest summary is that the API ports mechanically and the assumptions do not. Porting CUDA to HIP goes through exactly what to grep for.

which one to write in

SituationReasonable choice
Nvidia hardware only, and you want the newest features firstCUDA. New capabilities land there first and the ecosystem is deeper.
AMD hardware, now or laterHIP. It is the native interface and there is no reason to add a translation step.
You want one source tree for bothHIP, which compiles for both vendors.
You are writing kernels through a frameworkOften neither directly. PyTorch, Triton and the vendor libraries cover a great deal of ground before you need to write a kernel at all.

the rest of this section

Each of the following was a section of this page and is now a page of its own, with room for the code and the failure modes that a summary could not carry.

Your First GPU Kernel — vector addition in both dialects, complete, with the five host-side steps.
Thread Indexing and Grid-Stride Loops — the global index formula, the mandatory bounds check, and choosing block sizes.
Compiling GPU Code — nvcc against hipcc, PTX against SASS, and why AMD has no JIT fallback.
Porting CUDA to HIP — what hipify translates, and the four categories it cannot.
Streams — overlapping copies with computation, and the pinned memory requirement people miss.
Driver, Toolkit and Runtime — why nvidia-smi and nvcc report different CUDA versions.

related topics

CPU vs GPU: Why GPUs Exist — the design bet the whole programming model is built around.
What a Kernel Launch Physically Does — what the hardware does with a launch, as opposed to how you write one.
GPU Libraries — cuBLAS, rocBLAS, cuDNN, MIOpen and the rest, built on top of all this.
PyTorch Notes — how PyTorch uses CUDA underneath.

reference

NVIDIA CUDA C++ Programming Guide
AMD HIP documentation
AMD ROCm documentation
CUDA Runtime API reference