the short version

"What CUDA version do you have" is an ambiguous question, and the ambiguity is the source of an enormous amount of wasted time. There are three separate things with version numbers: the kernel-mode driver, the toolkit you compile with, and the runtime your program links against. They are deliberately allowed to differ.

The single most reported symptom of this confusion is that nvidia-smi and nvcc --version print different numbers. That is expected. They are not measuring the same thing.

the three things

ThingWhat it isHow you get itHow you check it
DriverKernel module and user-space libraries that talk to the hardwareInstalled on the host, usually by the system package managernvidia-smi, left-hand version
ToolkitThe compiler and development libraries, nvcc and friendsInstalled per project, or baked into a container imagenvcc --version
RuntimeThe library your binary actually calls at execution timeLinked in from the toolkit, or loaded from the systemReported by the API from inside your program

The version nvidia-smi shows in its top right is not the toolkit you have installed. It is the highest CUDA runtime version that your driver is capable of supporting. It answers "what could run here", not "what is installed here", which is exactly why it is usually the larger of the two numbers.

why they are allowed to differ

Drivers are a system-wide, privileged, disruptive thing to change. Toolkits are a per-project thing that different applications legitimately disagree about. Forcing them to match would mean every application on a machine had to agree on a CUDA version, and updating one would mean updating the driver underneath all the others.

So the contract runs one way: a newer driver supports applications built against older toolkits. Build against an older toolkit and run on a newer driver, and you are in the supported case. The reverse, building against a toolkit newer than the driver can support, is the case that fails, and it fails at load time with an error about an insufficient driver version rather than anything that names the real problem.

Within a major version, the rules are looser still: an application built with one minor version of the toolkit generally runs on a driver from any other minor version of that same major release, which is what makes it practical to ship binaries at all.

which one is your problem

SymptomUsually means
"CUDA driver version is insufficient for CUDA runtime version"Your toolkit is newer than the driver supports. Upgrade the driver, or build against an older toolkit.
"no CUDA-capable device is detected" with a card plainly presentA driver problem, not a toolkit problem. Often a kernel update that the module was not rebuilt against.
nvcc missing entirely, but nvidia-smi works fineYou have a driver and no toolkit. Very common on cloud images and inside containers.
Works on the host, fails in the containerThe container carries its own toolkit but must borrow the host driver. That is the container runtime's job to arrange.
A framework complains about a compute capabilityNeither of the three. The binary was built without support for your particular architecture.

containers, which is where this usually bites

The convention for GPU containers is that the image carries the toolkit and any CUDA libraries, while the driver stays on the host and is injected into the container at run time. This is exactly the one-way compatibility rule being used on purpose: the host driver is typically newer than whatever toolkit the image was built with, which is the supported direction.

The practical consequence is that a GPU container is not portable to a host whose driver is too old for it, and that "it works on my machine" has an extra dimension. When an image fails on one cluster and works on another, comparing host driver versions is the first thing to do.

the same picture on AMD

ROCm has the same structure with different names. There is a kernel driver, amdgpu, usually shipped with the distribution kernel or installed alongside ROCm. There is the ROCm stack itself, which supplies hipcc and the libraries and is the closest analogue to the toolkit. And there is a third number CUDA has no direct equivalent for: the gfx target, an identifier like gfx90a or gfx942 naming the exact GPU architecture your code is compiled for.

That last one is stricter than it looks. A binary built for one gfx target does not run on another, and ROCm supports a specific published list of targets per release, so "my card is AMD and ROCm is installed" is not sufficient. Check that your card's gfx target is on the supported list for the ROCm version you have, because an unsupported target is the single most common cause of a ROCm installation that appears complete and does nothing.

see it on your own machine

Run all of these. The point of the exercise is seeing them disagree.

nvidia-smi | https://developer.nvidia.com/nvidia-system-management-interface | driver version on the left, maximum supported CUDA runtime on the right. Not your toolkit. |'dtr_smi'
nvcc --version | https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html | the toolkit actually installed, which is the number that matters when you compile |'dtr_nvcc'
rocminfo | https://rocm.docs.amd.com/projects/rocminfo/en/latest/ | agent list with the gfx target of each AMD device, the number to check against the support list |'dtr_rocminfo'
hipcc --version | https://rocm.docs.amd.com/projects/HIP/en/latest/ | the ROCm compiler version, the AMD counterpart to nvcc --version |'dtr_hipcc'

related topics

CUDA & HIP — compiling with these toolchains once you know which one you have.
GPU History — why a separate toolkit and driver exist at all.
Consumer vs Datacenter GPUs — the other place where driver branches and licensing differ.

reference

NVIDIA CUDA C++ Programming Guide
NVIDIA CUDA compatibility
AMD ROCm documentation
AMD HIP documentation