Skip to main content
ZICQ

Wiki AI Concepts

Speculative Decoding

AI Concepts
Aliases: Speculative Decoding Speculative Sampling ·2026-09-19

Speculative Decoding

Speculative decoding is a popular LLM inference acceleration technique that emerged from 2023. Core idea:

Use a small, fast draft model to "guess" K tokens, then have the large model verify them in parallel. If the guesses are right, emit all K tokens at once; if not, fall back to the large model from the first mismatch. Since verification is parallel, overall latency approaches the draft model's latency while the output distribution is mathematically identical to the large model's.

How it works

  1. Draft model (e.g. 7B Llama) autoregressively generates K candidate tokens
  2. Target model (e.g. 70B Llama) one forward pass verifies all K tokens
  3. Accept the longest matching prefix (typically ≥ 2 tokens)
  4. From the first mismatch, the target model regenerates

Speedup

  • Typical speedup: 2-3× latency reduction (up to 4×)
  • Conditions: draft model distribution must be close to target
  • Small draft (100M-1B) + large target (7B-70B+) is the sweet spot

Key variants

Self-speculative decoding

  • No separate draft model — target model exits early (early exit), using intermediate layers as the draft
  • Saves memory but smaller speedup

Medusa / Lookahead decoding

  • Add lightweight "heads" to the target model to predict multiple future tokens directly
  • More memory-efficient than a separate draft model

Prompt-lookup decoding

  • Find repeated fragments in the prompt as candidate tokens
  • Good for "copy-style" tasks: code completion, document continuation

EAGLE / EAGLE-2

  • Train a lightweight drafter on the target model's top-layer hidden state
  • The current "best of breed" mainstream accelerator

Production adoption

  • vLLM, TensorRT-LLM, SGLang all built in
  • Major inference clouds (Fireworks / Together / Anyscale) enable by default
  • Composable with batch decoding, continuous batching, prefix caching

Strengths

  • Output distribution strictly preserved: mathematically proven that accept/reject sampling leaves the distribution unchanged
  • No retraining of the target model
  • Orthogonal to quantization (INT8 / FP8 / FP4); stacks with it

Limitations

  • Requires a good-enough draft model; bad draft → worse speedup
  • Larger speedup on long-form (code / book) than short-form (chat)
  • KV cache still must be managed; speedup is bounded by memory bandwidth