KV Cache
A cache of previously computed attention key and value tensors used to avoid recomputing earlier sequence states during autoregressive generation.
A KV cache stores key and value tensors from previous tokens during autoregressive Transformer inference so they do not have to be recomputed at every decoding step.
Language-Model Context
During autoregressive decoding, a KV cache stores attention key and value tensors computed for earlier tokens. Each new token can then reuse those states instead of recomputing the full prefix, substantially reducing repeated decoding work.
Memory Boundary
The KV cache is not a cache of model weights. Its memory footprint grows with sequence length, batch or concurrency, layer count, KV-head configuration, and datatype, so long context can make cache capacity a primary serving constraint.
Related LLM Concepts
Approximate Memory Formula
Exact size is architecture-dependent, but a useful lower-level capacity estimate can be written explicitly for conventional causal attention. Let L be layer count, H_kv the number of KV heads, T cached tokens, D_h head dimension, B the number of concurrent sequences or batch entries, and s bytes per stored element. Because both K and V are retained, tensor data alone scales approximately as:
M_KV ≈ 2 × L × H_kv × T × D_h × B × sThis estimate excludes allocator metadata, padding, fragmentation, and other runtime buffers, so it is a capacity approximation rather than an exact process-memory prediction. Grouped-Query Attention and MQA can reduce cache size by lowering H_kv. Reuse of past K/V tensors during autoregressive inference is documented directly in the Hugging Face Transformers cache documentation.
Related technical article: Artificial Intelligence: Philosophy, Theory and Practice.