queued next

  • Host vs device code — what __global__, __device__ and __host__ actually mean.
  • Catching silent GPU errors — proper error checking, and why a broken kernel often says nothing.
  • GPU memory management — malloc and memcpy against unified memory.
  • Pinned host memory — why page-locked transfers are faster.
  • Timing GPU code correctly — events, warmup, and the synchronize everyone forgets.
  • Writing a reduction kernel — summing an array the hard way, and why it is hard.
  • Warp-level primitives — shuffle, ballot and cooperative groups.
  • Writing a Triton kernel from scratch — starting with softmax.

already published in this section

CUDA & HIP
GPU Libraries
PyTorch on GPU
Your First GPU Kernel
Thread Indexing & Grid-Stride Loops
Compiling GPU Code
Porting CUDA to HIP
Streams & Async Work

the other GPU sections

GPU Hardware — what is queued there.
GPU Optimization — what is queued there.