Continuous Batching

Turkish equivalent: Sürekli batchingDomain: Machine Learning Systems

An inference-scheduling technique that continuously admits and retires generation requests so accelerator batches stay utilized despite variable sequence lengths.

ML-Systems Context

Continuous batching keeps accelerator batches populated by admitting new generation requests and retiring completed ones while other sequences are still decoding. This reduces the wasted work of static batches when request lengths differ substantially.

Service Boundary

The scheduler trades throughput against time to first token, inter-token latency, and fairness. Larger effective batches do not guarantee lower latency because queueing delay and KV-cache capacity can become the limiting resources.

Direct source: The primary paper or official specification for Continuous Batching is linked here for verification.