Paged Attention
An attention-memory management technique that stores KV-cache blocks in page-like units so variable-length requests can share accelerator memory more efficiently.
Related Concepts
- KV Cache
- Continuous Batching
- Memory Bandwidth
- Admission Control