CUDA & HIP

Kernel launch syntax, thread hierarchy, compiling with nvcc/hipcc, streams, and writing portable kernels with HIP.

GPU Libraries

BLAS (cuBLAS/rocBLAS/hipBLASLt), deep learning primitives (cuDNN/MIOpen), collectives (NCCL/RCCL), and kernel-generation frameworks (CUTLASS/Composable Kernel/Triton).

PyTorch on GPU

Mixed precision/AMP, torch.compile, the PyTorch profiler, the caching allocator, and multi-GPU training basics.

More Topics (Coming Soon)

More on this hub soon — your first kernel in CUDA and HIP side by side, thread indexing, error checking, porting with hipify, reduction kernels, and warp-level primitives. For the hardware these kernels run on see Anatomy of a GPU; for making them fast see GPU Optimization.

related topics

CPU vs GPU: Why GPUs Exist — why the hardware looks the way it does.
Anatomy of a GPU — the die, memory and host link these kernels run on.
What a Kernel Launch Physically Does — what happens between your launch call and work starting on the device.
SM vs CU — the compute block your thread blocks are assigned to.
Warps vs Wavefronts — the 32 or 64 thread group that is the real unit of execution.
The GPU Memory Hierarchy — registers to HBM, and what each level costs.
Tensor Cores vs Matrix Cores — the units that supply most of the chip's FLOPs.
FLOPs and TFLOPS — where the spec-sheet number comes from and why you never reach it.
Consumer vs Datacenter GPUs — what the price gap actually buys.
GPU History — why the programming model looks the way it does.
Driver, Toolkit and Runtime — what a CUDA version number actually refers to.
GPU Optimization — profiling and the techniques that make a working kernel fast.
Programming — the Python/PyTorch layer most GPU code is written from.
Developer Tools: Debugging & Profiling — debugging and profiling notes that apply directly to GPU code.
Machine Learning — the ML workloads this hardware and code exist to run.

reference

NVIDIA CUDA C++ Programming Guide
AMD ROCm documentation
GPU MODE (lectures on GPU performance)