KV Cache
KV Cache (Key-Value Cache) is the key optimization during inference: cache the Key and Value matrices of already-processed tokens so we don't recompute the entire history for every new token.
How it works
During Transformer decoding, generating each new token requires attention:
- Without KV cache: recompute K/V for full history each step, O(n²) complexity.
- With KV cache: only compute Q for the new token, reuse historical K/V, O(n) complexity.
Why it's a memory bottleneck
KV cache size scales with "layers × heads × seq_len × head_dim". A 7B model running 100k context can consume 20+ GB of memory on KV cache alone.
Optimization techniques
- PagedAttention (vLLM): page-store KV cache, eliminate fragmentation.
- GQA / MQA: multiple Queries share K/V heads (Llama 2/3 use this).
- Sliding Window Attention: only keep KV for the most recent N tokens (Mistral).
- KV cache quantization: quantize KV to INT8/INT4 too.
- FlashAttention: compute attention in SRAM, reduce HBM read/write.