vLLM
vLLM is the open-source high-throughput LLM inference engine from UC Berkeley RISELab. The PagedAttention paper made it the de facto industry standard.
Core innovation: PagedAttention
Traditional inference engines store KV cache in contiguous GPU memory, wasting a lot (internal + external fragmentation). vLLM splits KV cache into fixed-size "pages" (analogous to OS virtual memory):
- Pages from different requests can be mixed, near-zero fragmentation.
- Supports prefix sharing: requests with identical system prompts share KV cache pages.
- Memory utilization up 2-4x, throughput up 2-24x.
Main features
- Continuous batching: dynamically pack new requests into in-flight batches.
- Multiple quantizations: GPTQ, AWQ, INT4/INT8, FP8.
- Multi-GPU: tensor parallel / pipeline parallel.
- Multi-model: load any LLM directly from HuggingFace.
- OpenAI-compatible API:
/v1/chat/completionsdrops into existing tools.
Deployment
pip install vllm
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--gpu-memory-utilization 0.9
One line launches an OpenAI-compatible API.
Best for
- Production-grade API service (throughput first).
- Self-hosting as OpenAI replacement.
- Multi-model concurrency (routing + scheduling).