queued next

  • Vectorized loads — float4 and why wide beats narrow.
  • Why FlashAttention exists — optimizing a memory-bound softmax.

already published in this section

GPU Optimization Overview
Assembly & Low-Level GPU Systems
The Optimization Mindset
GPU Profiling Tools
The Roofline Model
Memory Coalescing
Occupancy & Register Pressure
Shared Memory & Bank Conflicts
Kernel Fusion
Mixed Precision
Tiling & Blocking for Reuse
Autotuning & Kernel Generation

the other GPU sections

GPU Hardware — what is queued there.
GPU Programming — what is queued there.