Paged Attention

Turkish equivalent: Sayfalı attentionDomain: Machine Learning Systems

An attention-memory management technique that stores KV-cache blocks in page-like units so variable-length requests can share accelerator memory more efficiently.

  • KV Cache
  • Continuous Batching
  • Memory Bandwidth
  • Admission Control