Autotuning and Kernel Generation
Past a certain point the right tile size is not something you reason about. It is something you search for.
A kernel has a handful of structural parameters: tile dimensions, how much each thread computes, unroll factors, pipeline depth. The best combination depends on the problem shape, the architecture, the precision, and the interaction between all of them. It is not stable across any of those, and it is not something intuition predicts well.
So the modern approach stops guessing. Express the kernel as a template over those parameters, generate many concrete versions, benchmark them on the actual hardware, and keep whichever wins.
The parameters interact. A larger tile improves reuse but costs shared memory, which reduces resident blocks, which reduces latency hiding, which may or may not matter depending on how memory-bound the kernel was to begin with. Each parameter's effect depends on the others, so the space is not separable and tuning one at a time finds local optima.
And the optimum moves. A configuration tuned for a square matrix is wrong for a tall thin one. A configuration tuned for one architecture is wrong for the next, because register file sizes, scratchpad capacity and matrix unit shapes all changed. Anything hand-tuned is tuned for a moment.
The pattern is the same across the tools: declare the search space, let the framework instantiate and time each candidate, cache the winner keyed by the input shape.
@triton.autotune(
configs=[
triton.Config({'BLOCK_M': 128, 'BLOCK_N': 128, 'BLOCK_K': 32}, num_warps=8),
triton.Config({'BLOCK_M': 64, 'BLOCK_N': 128, 'BLOCK_K': 32}, num_warps=4),
triton.Config({'BLOCK_M': 128, 'BLOCK_N': 64, 'BLOCK_K': 64}, num_warps=4),
],
key=['M', 'N', 'K'],
)
@triton.jit
def matmul_kernel(...):
...
key is what makes the cache useful: results are stored per problem shape, so a new shape triggers a fresh search while a repeated shape reuses the previous winner. Get the key wrong, either too coarse or too fine, and you either use a badly matched configuration or re-tune constantly.
This is where autotuning becomes visible from outside. The first call with a new shape compiles and benchmarks a set of candidates, which can take from milliseconds to a noticeable pause. Subsequent calls with that shape hit the cache and are fast.
The same mechanism explains several familiar behaviours: why a compiled PyTorch model is slow on its first batch and fast afterwards, why the convolution benchmark flag helps steady-state training and hurts a script that runs once, and why variable sequence lengths can make a served model behave erratically as each new shape pays its own search.
The mitigations are the obvious ones: warm up with representative shapes before measuring or serving, bucket or pad shapes so that fewer distinct keys occur, and persist the tuning cache across runs where the framework supports it.
| Tool | What it generates | Where the tuning lives |
|---|---|---|
| Vendor BLAS | Precompiled kernels | Tuned by the vendor, selected by heuristic at call time |
| CUTLASS / Composable Kernel | C++ templates over tile shapes | You instantiate and benchmark, or use their profiler |
| Triton | A kernel from Python, compiled per configuration | The decorator above, cached per shape |
| torch.compile | Fused kernels for a whole graph | Automatic, with its own cache |
Moving down that list trades control for effort. The library is fastest to use and cannot express a fusion it does not already have; the template layer can express almost anything and asks you to search for it.
When the kernel is memory-bound and already near its bandwidth roof. No tile size recovers bandwidth that does not exist. Fix the access pattern or the traffic first.
When shapes are unstable. If every call has a new shape, you pay search cost constantly and never amortize it. Bucket the shapes or use a fixed reasonable configuration instead.
When the search space is wrong. Autotuning finds the best of what you offered it. If every candidate shares a flawed structure, it returns the least bad version of that flaw, quite convincingly.