the short version

There are two distinct questions and they need different tools. "Where is my program's time going, and is the GPU even busy?" is a timeline question. "Why is this one kernel slow?" is a kernel question. A timeline profiler cannot tell you about bank conflicts; a kernel profiler cannot tell you that you left the device idle for 200 milliseconds waiting on a data loader.

Ask the first question before the second. Optimizing a kernel that accounts for 3% of runtime is a very common way to achieve nothing.

the tools

QuestionNvidiaAMDPyTorch
Is the device busy at all?nvidia-smirocm-smi—
Where does wall clock go?Nsight Systemsrocprofv3, omnitracePyTorch Profiler
Why is this kernel slow?Nsight Computerocprofv3 counters—
Which op in my model?——PyTorch Profiler

start with the dumbest check

Before installing anything, watch utilisation while the job runs. If it sits near zero, the problem is not in any kernel: you are starved by the input pipeline, blocked on host work, or synchronizing constantly. No kernel optimization fixes that.

nvidia-smi --query-gpu=utilization.gpu,memory.used,power.draw --format=csv --loop-ms=200 | https://developer.nvidia.com/nvidia-system-management-interface | live utilisation, memory and power; the fastest way to learn the device is idle half the time |'pf_smi'
rocm-smi --showuse --showmemuse | https://rocm.docs.amd.com/projects/rocm_smi_lib/en/latest/ | the same first check on AMD hardware |'pf_rocmsmi'

Power draw is a useful secondary signal. A device at high utilisation but low power is often memory-bound and stalling rather than doing arithmetic.

timeline profilers

These record what happened when, across host and device, and show it as tracks you can line up. What you are looking for is gaps: device idle while the host prepares work, kernels that could overlap but do not, copies that serialize against compute, and unexpected synchronization points.

nsys profile -o report --stats=true ./my_app | https://docs.nvidia.com/nsight-systems/ | the standard first profile on Nvidia; open the report in the GUI for the timeline, or read the summary tables in the terminal |'pf_nsys'
rocprofv3 --kernel-trace --hip-trace --stats -- ./my_app | https://rocm.docs.amd.com/projects/rocprofiler-sdk/en/latest/ | device activity alongside HIP API calls, the AMD equivalent view |'pf_rocprofv3'

The output also gives you the ranked list of kernels by total time, which is what tells you which kernel is worth the next section.

kernel profilers

These replay a single kernel with hardware counters enabled and report what it was limited by: achieved bandwidth, achieved arithmetic throughput, occupancy, cache hit rates, bank conflicts, and the fraction of lanes actually doing work.

ncu --set full -o profile ./my_app | https://docs.nvidia.com/nsight-compute/ | the full report, including a roofline chart and explicit warnings about uncoalesced access and low occupancy |'pf_ncu'
ncu --kernel-name myKernel --launch-count 1 --set full ./my_app | https://docs.nvidia.com/nsight-compute/ | profile one kernel and one launch; a full profile of everything is slow because each kernel is replayed |'pf_ncu_one'

Be aware that kernel profiling is invasive. The tool replays kernels to collect different counter sets, so wall clock under the profiler is not your real runtime. Use it for ratios and limiters, not for timing.

profiling from inside the framework

When the question is "which layer of my model is slow", a framework-level profiler maps device time back to the operations you wrote, which neither of the tools above can do.


from torch.profiler import profile, ProfilerActivity

with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
             record_shapes=True) as prof:
    model(inputs)

print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=15))
            
This gives a ranked table of operations by device time, with shapes, which is usually enough to find the layer worth attacking. It also exports a trace you can open in a timeline viewer when you need to see the overlap.

a workflow

Check utilisation. If it is low, fix the pipeline and stop. If it is high, take a timeline profile and find the kernel that dominates. Profile that kernel and find its limiter. Change one thing, measure again outside the profiler, and repeat.

Keep the before and after numbers. A change that is obviously an improvement often is not, and the only defence against remembering it as a win is having written the number down.

related topics

The Optimization Mindset — what to do with the numbers these tools give you.
The Roofline Model — the chart Nsight Compute draws, explained.
Occupancy and Register Pressure — one of the headline numbers in every kernel report.

reference

NVIDIA Nsight Systems
NVIDIA Nsight Compute
AMD rocprofiler
PyTorch Profiler