Prefix Caching
Reuse of previously computed model state for a shared prompt prefix so repeated requests avoid recomputing identical early tokens.
ML-Systems Context
Prefix caching reuses previously computed model state when multiple requests begin with the same token prefix. Workloads with a long shared system prompt or common context can reduce time to first token and accelerator compute by avoiding repeated processing of those early tokens.
Cache Boundary
Semantically similar prompts are not necessarily the same prefix. Cache identity must safely include the exact token sequence and relevant model, tokenizer, adapter, and version context.