KV Cache
A cache of previously computed attention key and value tensors used to avoid recomputing earlier sequence states during autoregressive generation.
Related Concepts
- Context Window
- Multi-Head Attention
- Speculative Decoding
- Quantization
A cache of previously computed attention key and value tensors used to avoid recomputing earlier sequence states during autoregressive generation.