Unlocking Asynchronicity in Continuous Batching

Hugging Face shows how asynchronous continuous batching, using CUDA streams, events, and double-buffered slots, lifted GPU utilization from 76 to 99.4 percent and cut generation time by 22 percent.

Thursday May 14, 2026 Source: huggingface.co
TL;DR — Quick Answer

Asynchronous continuous batching separates CPU batch preparation from GPU compute so both run in parallel instead of taking turns. In Hugging Face's published experiments, generating 8K tokens with an 8B model at batch size 32 took 300.6 seconds with synchronous batching and dropped to 234.5 seconds with the async loop — a 22 percent speedup with no new kernels or model changes. GPU utilization rose from 76.0 percent to 99.4 percent. The technique uses three CUDA streams for transfers and compute, CUDA events to enforce ordering, double-buffered slots to avoid race conditions, a memory pool for CUDA graphs, and a carry-over mask to move batch N output tokens into batch N+1 inputs. The implementation is live in the Hugging Face transformers library.

Key Takeaways

Unlocking Asynchronicity in Continuous Batching — AI news article illustration

An idle GPU is expensive. Hugging Face’s second post in its efficient LLM inference series tackles the wasted cycles that continuous batching alone does not fix: default synchronous scheduling, where the CPU and GPU take turns instead of working together. The result is asynchronous batching, and the payoff is a substantial speedup with zero model changes.

The Sync Bottleneck

Continuous batching packs requests tightly to avoid padding waste, but the loop runs synchronously by default. While the GPU computes a batch, the CPU waits; while the CPU re-schedules and prepares the next batch, the GPU waits. In a loop running hundreds of steps per second, those idle gaps add up.

Profiling showed the scale of the problem: generating 8K tokens with an 8B model at batch size 32 took 300.6 seconds, with 24.0 percent of that time spent with an idle GPU waiting on the CPU.

CUDA Streams and Events

The fix disentangles CPU prep from GPU compute. Operations are split across three CUDA streams:

Since streams are independent, ordering is enforced with CUDA events: record marks when a transfer finishes and wait blocks a downstream stream until then. The default stream is avoided because it is synchronizing and would defeat the concurrency.

Double Buffering and Carry-Over

Two problems remain. First, reusing the same tensors for batch N and N+1 would corrupt data the GPU is still reading — a race condition. The answer is double-buffered input and output slots, allowed to share one memory pool so both CUDA graphs use nearly the same VRAM as a single graph.

Second, a request in both batches produces a token in batch N that is its input in N+1. Since that token is not ready when N+1 is prepared, a carry-over mask and placeholder tokens let the GPU patch batch N outputs into batch N+1 inputs inside the captured graph.

The Results

With the async loop, the GPU is active for 99.4 percent of total runtime, up from 76.0 percent. Total generation time fell from 300.6 seconds to 234.5 seconds — a 22 percent speedup, closing most of the theoretical 24 percent ceiling. No new kernels, no model changes — just letting the CPU and GPU work at the same time.

Why It Matters

The technique moves inference from schedule-based to data-based dependencies and is a stepping stone to state-of-the-art throughput for long generations of 16K tokens and beyond, as used in reinforcement learning. The full implementation ships in the Hugging Face transformers continuous batching module.

Frequently Asked Questions

What is async continuous batching?

Asynchronous continuous batching prepares the next batch on the CPU while the GPU is still computing the current one, using CUDA streams and events so both run in parallel and the GPU never idles between batches.

How much faster is asynchronous batching?

In Hugging Face's tests, generating 8K tokens with an 8B model at batch size 32 dropped from 300.6 to 234.5 seconds — a 22 percent speedup with no kernel or model changes.

What are CUDA streams and events used for here?

Three streams handle transfers and compute, and CUDA events tell the compute stream to wait for the input transfer and the output transfer to wait for compute, preserving order without blocking the CPU.

Why does the GPU go idle in synchronous batching?

In synchronous batching the CPU and GPU take turns: the GPU computes, then waits idle while the CPU re-schedules the batch and prepares the next input. In continuous loops this idling can consume nearly a quarter of total runtime.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news