An idle GPU is expensive. Hugging Face’s second post in its efficient LLM inference series tackles the wasted cycles that continuous batching alone does not fix: default synchronous scheduling, where the CPU and GPU take turns instead of working together. The result is asynchronous batching, and the payoff is a substantial speedup with zero model changes.
The Sync Bottleneck
Continuous batching packs requests tightly to avoid padding waste, but the loop runs synchronously by default. While the GPU computes a batch, the CPU waits; while the CPU re-schedules and prepares the next batch, the GPU waits. In a loop running hundreds of steps per second, those idle gaps add up.
Profiling showed the scale of the problem: generating 8K tokens with an 8B model at batch size 32 took 300.6 seconds, with 24.0 percent of that time spent with an idle GPU waiting on the CPU.
CUDA Streams and Events
The fix disentangles CPU prep from GPU compute. Operations are split across three CUDA streams:
- H2D stream — transfers input tensors from CPU to GPU
- Compute stream — runs the model forward pass
- D2H stream — copies outputs back to the CPU
Since streams are independent, ordering is enforced with CUDA events: record marks when a transfer finishes and wait blocks a downstream stream until then. The default stream is avoided because it is synchronizing and would defeat the concurrency.
Double Buffering and Carry-Over
Two problems remain. First, reusing the same tensors for batch N and N+1 would corrupt data the GPU is still reading — a race condition. The answer is double-buffered input and output slots, allowed to share one memory pool so both CUDA graphs use nearly the same VRAM as a single graph.
Second, a request in both batches produces a token in batch N that is its input in N+1. Since that token is not ready when N+1 is prepared, a carry-over mask and placeholder tokens let the GPU patch batch N outputs into batch N+1 inputs inside the captured graph.
- Doubled buffers prevent CPU writes from racing GPU reads
- A memory pool keeps CUDA graph VRAM near-constant
- Carry-over runs as four cheap tensor ops inside the graph
The Results
With the async loop, the GPU is active for 99.4 percent of total runtime, up from 76.0 percent. Total generation time fell from 300.6 seconds to 234.5 seconds — a 22 percent speedup, closing most of the theoretical 24 percent ceiling. No new kernels, no model changes — just letting the CPU and GPU work at the same time.
Why It Matters
The technique moves inference from schedule-based to data-based dependencies and is a stepping stone to state-of-the-art throughput for long generations of 16K tokens and beyond, as used in reinforcement learning. The full implementation ships in the Hugging Face transformers continuous batching module.