Paged Attention

Turkish equivalent: Sayfalı attentionDomain: Machine Learning Systems

An attention-memory management technique that stores KV-cache blocks in page-like units so variable-length requests can share accelerator memory more efficiently.

ML-Systems Context

Paged attention stores KV-cache state in page-like blocks rather than requiring one large contiguous allocation per request. This can improve memory utilization and request admission under variable sequence lengths and high concurrency, especially when paired with continuous batching.

Model Boundary

Paged attention primarily changes serving-time memory management. It does not inherently change the mathematical definition of the model's attention operation.

What the Virtual-Memory Analogy Means

The central idea of PagedAttention is not to redefine the attention equation, but to avoid requiring each request's KV Cache to occupy one large contiguous allocation. Logical KV blocks are mapped to physical blocks, reducing fragmentation and unnecessary copying for variable-length requests. The same abstraction can also support controlled sharing of some cached prefix state.

This is a serving-layer optimization. Using PagedAttention does not inherently change model logits or the mathematical meaning of attention. In the vLLM paper, Kwon et al. report 2–4× throughput improvements at a similar latency level under their evaluated workloads; that figure is an experimental result tied to the paper's models, hardware, and request mix rather than a universal multiplier. Source: PagedAttention paper.