#CUDA

2 articles about CUDA

Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP
AI Tools Jun 11, 2026

Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP

Part 2 of Hugging Face's PyTorch profiling series shows how to read torch.profiler traces from nn.Linear through compiled and hand-tuned MLP kernels, and what fuse decisions actually improve.

Unlocking Asynchronicity in Continuous Batching
AI Infrastructure May 14, 2026

Unlocking Asynchronicity in Continuous Batching

Hugging Face shows how asynchronous continuous batching, using CUDA streams, events, and double-buffered slots, lifted GPU utilization from 76 to 99.4 percent and cut generation time by 22 percent.