Driver, Toolkit and Runtime: What a CUDA Version Really Is
There are at least three of them, they are allowed to differ, and nvidia-smi and nvcc disagreeing is normal.
"What CUDA version do you have" is an ambiguous question, and the ambiguity is the source of an enormous amount of wasted time. There are three separate things with version numbers: the kernel-mode driver, the toolkit you compile with, and the runtime your program links against. They are deliberately allowed to differ.
The single most reported symptom of this confusion is that nvidia-smi and nvcc --version print different numbers. That is expected. They are not measuring the same thing.
| Thing | What it is | How you get it | How you check it |
|---|---|---|---|
| Driver | Kernel module and user-space libraries that talk to the hardware | Installed on the host, usually by the system package manager | nvidia-smi, left-hand version |
| Toolkit | The compiler and development libraries, nvcc and friends | Installed per project, or baked into a container image | nvcc --version |
| Runtime | The library your binary actually calls at execution time | Linked in from the toolkit, or loaded from the system | Reported by the API from inside your program |
The version nvidia-smi shows in its top right is not the toolkit you have installed. It is the highest CUDA runtime version that your driver is capable of supporting. It answers "what could run here", not "what is installed here", which is exactly why it is usually the larger of the two numbers.
Drivers are a system-wide, privileged, disruptive thing to change. Toolkits are a per-project thing that different applications legitimately disagree about. Forcing them to match would mean every application on a machine had to agree on a CUDA version, and updating one would mean updating the driver underneath all the others.
So the contract runs one way: a newer driver supports applications built against older toolkits. Build against an older toolkit and run on a newer driver, and you are in the supported case. The reverse, building against a toolkit newer than the driver can support, is the case that fails, and it fails at load time with an error about an insufficient driver version rather than anything that names the real problem.
Within a major version, the rules are looser still: an application built with one minor version of the toolkit generally runs on a driver from any other minor version of that same major release, which is what makes it practical to ship binaries at all.
| Symptom | Usually means |
|---|---|
| "CUDA driver version is insufficient for CUDA runtime version" | Your toolkit is newer than the driver supports. Upgrade the driver, or build against an older toolkit. |
| "no CUDA-capable device is detected" with a card plainly present | A driver problem, not a toolkit problem. Often a kernel update that the module was not rebuilt against. |
| nvcc missing entirely, but nvidia-smi works fine | You have a driver and no toolkit. Very common on cloud images and inside containers. |
| Works on the host, fails in the container | The container carries its own toolkit but must borrow the host driver. That is the container runtime's job to arrange. |
| A framework complains about a compute capability | Neither of the three. The binary was built without support for your particular architecture. |
The convention for GPU containers is that the image carries the toolkit and any CUDA libraries, while the driver stays on the host and is injected into the container at run time. This is exactly the one-way compatibility rule being used on purpose: the host driver is typically newer than whatever toolkit the image was built with, which is the supported direction.
The practical consequence is that a GPU container is not portable to a host whose driver is too old for it, and that "it works on my machine" has an extra dimension. When an image fails on one cluster and works on another, comparing host driver versions is the first thing to do.
ROCm has the same structure with different names. There is a kernel driver, amdgpu, usually shipped with the distribution kernel or installed alongside ROCm. There is the ROCm stack itself, which supplies hipcc and the libraries and is the closest analogue to the toolkit. And there is a third number CUDA has no direct equivalent for: the gfx target, an identifier like gfx90a or gfx942 naming the exact GPU architecture your code is compiled for.
That last one is stricter than it looks. A binary built for one gfx target does not run on another, and ROCm supports a specific published list of targets per release, so "my card is AMD and ROCm is installed" is not sufficient. Check that your card's gfx target is on the supported list for the ROCm version you have, because an unsupported target is the single most common cause of a ROCm installation that appears complete and does nothing.
Run all of these. The point of the exercise is seeing them disagree.