Compiling GPU Code: nvcc vs hipcc
One source file, two compilers, and an architecture flag that decides whether it runs at all.
A GPU source file contains code for two different processors. The compiler splits it, hands the host half to an ordinary C++ compiler, compiles the device half for one or more GPU architectures, and packs everything into a single executable. The flag that says which GPU architectures is the one that causes the most trouble.
Compiling with nvcc or hipcc gives you one binary containing host machine code plus one or more compiled copies of every kernel, each built for a particular device architecture. At run time the runtime picks the copy matching the device it finds.
If no copy matches, the two vendors behave very differently, and that difference is the single most important thing on this page.
Nvidia compiles through an intermediate representation. PTX is a virtual instruction set, forward compatible and not tied to a specific chip. SASS is the real machine code for one architecture. A binary can carry either or both.
| Flag form | What you get |
|---|---|
-arch=sm_90 | SASS for that architecture, and typically PTX too |
-gencode arch=compute_80,code=sm_80 | SASS for exactly sm_80 |
-gencode arch=compute_80,code=compute_80 | PTX only, to be JIT-compiled at run time |
Because PTX exists, a binary carrying it can run on a newer GPU than it was built for: the driver just-in-time compiles the PTX on first launch. That is a genuine safety net, at the cost of a pause on startup and code that is usually a little slower than if you had compiled for the target directly.
HIP compiles straight to a device ISA identified by a gfx target, such as gfx90a or gfx942. There is no widely used virtual ISA layer equivalent to PTX in the normal workflow, which means there is no JIT fallback to save you.
hipcc -O3 --offload-arch=gfx942 kernel.cpp -o kernel
# several targets in one binary
hipcc -O3 --offload-arch=gfx90a --offload-arch=gfx942 kernel.cpp -o kernel
# whatever is installed in this machine, handy while developing
hipcc -O3 --offload-arch=native kernel.cpp -o kernel
Shipping to more than one GPU means building for more than one target, and the binary grows with each. The common approach on Nvidia is to list the architectures you support plus PTX for the newest, so that future hardware still runs.
nvcc -O3 \
-gencode arch=compute_80,code=sm_80 \
-gencode arch=compute_90,code=sm_90 \
-gencode arch=compute_90,code=compute_90 \
kernel.cu -o kernel
| Symptom | Cause |
|---|---|
| Runs, but slower than expected, with a pause on first launch | No matching SASS, so the driver JIT-compiled the embedded PTX. |
| "no kernel image is available for execution on the device" | Nothing in the binary matches the device and there was no PTX to fall back on. The usual AMD failure, and the Nvidia one when PTX was omitted. |
| Compiles fine, launches fine, results are silently wrong | Not a compilation problem. Look for a missing bounds check or a race. |
| "unsupported gpu architecture" | The toolkit is older than the architecture you asked for. Upgrade the toolkit. |
| Host code compiles, device code does not, with template errors | The device compiler is stricter about what it accepts in device functions, especially around standard library use. |