2 articles about CUDA
Part 2 of Hugging Face's PyTorch profiling series shows how to read torch.profiler traces from nn.Linear through compiled and hand-tuned MLP kernels, and what fuse decisions actually improve.
Hugging Face shows how asynchronous continuous batching, using CUDA streams, events, and double-buffered slots, lifted GPU utilization from 76 to 99.4 percent and cut generation time by 22 percent.