the short version

A GPU source file contains code for two different processors. The compiler splits it, hands the host half to an ordinary C++ compiler, compiles the device half for one or more GPU architectures, and packs everything into a single executable. The flag that says which GPU architectures is the one that causes the most trouble.

what the compiler produces

Compiling with nvcc or hipcc gives you one binary containing host machine code plus one or more compiled copies of every kernel, each built for a particular device architecture. At run time the runtime picks the copy matching the device it finds.

If no copy matches, the two vendors behave very differently, and that difference is the single most important thing on this page.

Nvidia: two levels, and a fallback

Nvidia compiles through an intermediate representation. PTX is a virtual instruction set, forward compatible and not tied to a specific chip. SASS is the real machine code for one architecture. A binary can carry either or both.

Flag formWhat you get
-arch=sm_90SASS for that architecture, and typically PTX too
-gencode arch=compute_80,code=sm_80SASS for exactly sm_80
-gencode arch=compute_80,code=compute_80PTX only, to be JIT-compiled at run time

Because PTX exists, a binary carrying it can run on a newer GPU than it was built for: the driver just-in-time compiles the PTX on first launch. That is a genuine safety net, at the cost of a pause on startup and code that is usually a little slower than if you had compiled for the target directly.

AMD: one level, and no fallback

HIP compiles straight to a device ISA identified by a gfx target, such as gfx90a or gfx942. There is no widely used virtual ISA layer equivalent to PTX in the normal workflow, which means there is no JIT fallback to save you.


hipcc -O3 --offload-arch=gfx942 kernel.cpp -o kernel

# several targets in one binary
hipcc -O3 --offload-arch=gfx90a --offload-arch=gfx942 kernel.cpp -o kernel

# whatever is installed in this machine, handy while developing
hipcc -O3 --offload-arch=native kernel.cpp -o kernel
            
The practical rule is that you must name every architecture you intend to run on, at build time. A binary built for gfx90a will not run on gfx942, and the failure happens at launch rather than at load, often with a message about no available kernel image.

targeting several architectures

Shipping to more than one GPU means building for more than one target, and the binary grows with each. The common approach on Nvidia is to list the architectures you support plus PTX for the newest, so that future hardware still runs.


nvcc -O3 \
  -gencode arch=compute_80,code=sm_80 \
  -gencode arch=compute_90,code=sm_90 \
  -gencode arch=compute_90,code=compute_90 \
  kernel.cu -o kernel
            
The first two lines give native code for Ampere and Hopper. The third embeds PTX, so an architecture newer than Hopper can still JIT rather than fail. This is why a CUDA library binary can be surprisingly large.

what each mistake looks like

SymptomCause
Runs, but slower than expected, with a pause on first launchNo matching SASS, so the driver JIT-compiled the embedded PTX.
"no kernel image is available for execution on the device"Nothing in the binary matches the device and there was no PTX to fall back on. The usual AMD failure, and the Nvidia one when PTX was omitted.
Compiles fine, launches fine, results are silently wrongNot a compilation problem. Look for a missing bounds check or a race.
"unsupported gpu architecture"The toolkit is older than the architecture you asked for. Upgrade the toolkit.
Host code compiles, device code does not, with template errorsThe device compiler is stricter about what it accepts in device functions, especially around standard library use.

a few flags worth knowing

nvcc --ptxas-options=-v -O3 -arch=sm_90 kernel.cu -o kernel | https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/ | prints registers per thread and shared memory per block, the numbers that decide occupancy |'cg_ptxasv'
nvcc --list-gpu-arch | https://docs.nvidia.com/cuda/cuda-compiler-driver-nvcc/ | the architectures this toolkit knows how to build for |'cg_listarch'
hipcc --offload-arch=native -O3 kernel.cpp -o kernel | https://rocm.docs.amd.com/projects/HIP/en/latest/ | builds for the installed device without you having to look its gfx target up |'cg_native'
rocminfo | https://rocm.docs.amd.com/projects/rocminfo/en/latest/ | when you do need to look it up: the gfx target is listed per agent |'cg_rocminfo'

related topics

Driver, Toolkit and Runtime — which CUDA version the toolkit is, and why it differs from the driver.
Your First GPU Kernel — the program these commands are building.
Porting CUDA to HIP — building one source tree for both toolchains.

reference

NVIDIA nvcc compiler driver
AMD HIP documentation
NVIDIA CUDA C++ Programming Guide