GPU Profiling Tools
Two different questions, two different tools. Reaching for the wrong one wastes an afternoon.
There are two distinct questions and they need different tools. "Where is my program's time going, and is the GPU even busy?" is a timeline question. "Why is this one kernel slow?" is a kernel question. A timeline profiler cannot tell you about bank conflicts; a kernel profiler cannot tell you that you left the device idle for 200 milliseconds waiting on a data loader.
Ask the first question before the second. Optimizing a kernel that accounts for 3% of runtime is a very common way to achieve nothing.
| Question | Nvidia | AMD | PyTorch |
|---|---|---|---|
| Is the device busy at all? | nvidia-smi | rocm-smi | — |
| Where does wall clock go? | Nsight Systems | rocprofv3, omnitrace | PyTorch Profiler |
| Why is this kernel slow? | Nsight Compute | rocprofv3 counters | — |
| Which op in my model? | — | — | PyTorch Profiler |
Before installing anything, watch utilisation while the job runs. If it sits near zero, the problem is not in any kernel: you are starved by the input pipeline, blocked on host work, or synchronizing constantly. No kernel optimization fixes that.
Power draw is a useful secondary signal. A device at high utilisation but low power is often memory-bound and stalling rather than doing arithmetic.
These record what happened when, across host and device, and show it as tracks you can line up. What you are looking for is gaps: device idle while the host prepares work, kernels that could overlap but do not, copies that serialize against compute, and unexpected synchronization points.
The output also gives you the ranked list of kernels by total time, which is what tells you which kernel is worth the next section.
These replay a single kernel with hardware counters enabled and report what it was limited by: achieved bandwidth, achieved arithmetic throughput, occupancy, cache hit rates, bank conflicts, and the fraction of lanes actually doing work.
Be aware that kernel profiling is invasive. The tool replays kernels to collect different counter sets, so wall clock under the profiler is not your real runtime. Use it for ratios and limiters, not for timing.
When the question is "which layer of my model is slow", a framework-level profiler maps device time back to the operations you wrote, which neither of the tools above can do.
from torch.profiler import profile, ProfilerActivity
with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
record_shapes=True) as prof:
model(inputs)
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=15))
Check utilisation. If it is low, fix the pipeline and stop. If it is high, take a timeline profile and find the kernel that dominates. Profile that kernel and find its limiter. Change one thing, measure again outside the profiler, and repeat.
Keep the before and after numbers. A change that is obviously an improvement often is not, and the only defence against remembering it as a win is having written the number down.