Prefix Caching
Reuse of previously computed model state for a shared prompt prefix so repeated requests avoid recomputing identical early tokens.
Related Concepts
- KV Cache
- Context Window
- Paged Attention
- Cache Stampede
Reuse of previously computed model state for a shared prompt prefix so repeated requests avoid recomputing identical early tokens.