Skip to main content
ZICQ

Wiki Infrastructure

vLLM

Infrastructure
Aliases: vLLM ·2026-09-14

vLLM

vLLM is the open-source high-throughput LLM inference engine from UC Berkeley RISELab. The PagedAttention paper made it the de facto industry standard.

Core innovation: PagedAttention

Traditional inference engines store KV cache in contiguous GPU memory, wasting a lot (internal + external fragmentation). vLLM splits KV cache into fixed-size "pages" (analogous to OS virtual memory):

  • Pages from different requests can be mixed, near-zero fragmentation.
  • Supports prefix sharing: requests with identical system prompts share KV cache pages.
  • Memory utilization up 2-4x, throughput up 2-24x.

Main features

  • Continuous batching: dynamically pack new requests into in-flight batches.
  • Multiple quantizations: GPTQ, AWQ, INT4/INT8, FP8.
  • Multi-GPU: tensor parallel / pipeline parallel.
  • Multi-model: load any LLM directly from HuggingFace.
  • OpenAI-compatible API: /v1/chat/completions drops into existing tools.

Deployment

pip install vllm
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3-8B-Instruct \
    --gpu-memory-utilization 0.9

One line launches an OpenAI-compatible API.

Best for

  • Production-grade API service (throughput first).
  • Self-hosting as OpenAI replacement.
  • Multi-model concurrency (routing + scheduling).