Skip to main content
ZICQ

Wiki Concepts

KV Cache

Concepts
Aliases: Key-Value Cache ·2026-09-14

KV Cache

KV Cache (Key-Value Cache) is the key optimization during inference: cache the Key and Value matrices of already-processed tokens so we don't recompute the entire history for every new token.

How it works

During Transformer decoding, generating each new token requires attention:

  • Without KV cache: recompute K/V for full history each step, O(n²) complexity.
  • With KV cache: only compute Q for the new token, reuse historical K/V, O(n) complexity.

Why it's a memory bottleneck

KV cache size scales with "layers × heads × seq_len × head_dim". A 7B model running 100k context can consume 20+ GB of memory on KV cache alone.

Optimization techniques

  • PagedAttention (vLLM): page-store KV cache, eliminate fragmentation.
  • GQA / MQA: multiple Queries share K/V heads (Llama 2/3 use this).
  • Sliding Window Attention: only keep KV for the most recent N tokens (Mistral).
  • KV cache quantization: quantize KV to INT8/INT4 too.
  • FlashAttention: compute attention in SRAM, reduce HBM read/write.