Docker
Docker is the de facto standard containerization platform. Packages applications
- dependencies + config into images, build once, run anywhere. Mainstream form for LLM service deployment (vllm / ollama / microservices).
Why LLM projects all use Docker
- Environment consistency: same image across dev / test / prod, avoid "works on my machine".
- GPU passthrough:
--gpus alldirectly uses host NVIDIA driver, no driver installation in container. - Model weights mount: large model weights via volume mount, images stay slim.
- Fast replication: spin up multiple dev / staging / prod instances with one command.
Core concepts
- Image: read-only template (layered structure, shared base layers).
- Container: running instance of an image.
- Dockerfile: build script for an image.
- Volume: persistent data / config files mount.
- Network: container-to-container network (bridge / host / overlay).
- Registry: image repository (Docker Hub, GHCR, private Harbor / ECR).
Typical Dockerfile for LLM service
FROM nvcr.io/nvidia/pytorch:24.05-py3 # Pre-installed CUDA + PyTorch
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
EXPOSE 8000
CMD ["python", "-m", "vllm.entrypoints.openai.api_server", \
"--model", "/models/llama-3.1-70b", \
"--tensor-parallel-size", "2"]
docker-compose multi-service orchestration
services:
vllm:
image: vllm/vllm-openai:latest
runtime: nvidia # or deploy.resources.reservations.devices
volumes:
- /data/models:/models
ports:
- "8000:8000"
app:
build: .
depends_on: [vllm, postgres]
environment:
VLLM_BASE_URL: http://vllm:8000
Practical experience
- Don't install NVIDIA driver in container: use
--gpuspassthrough, CUDA toolkit in image, driver on host. - Multi-stage build for smaller image: build stage installs dev deps, runtime stage only COPYs build output.
- Run as non-root user:
USER 10001, avoid container escape attacks. - Read-only rootfs: container root partition read-only, config / data via volume, minimize attack surface.
- Health checks: each service configure
HEALTHCHECK, compose / k8s determines readiness from it.
Docker vs Podman / nerdctl
- Docker: most mature ecosystem, all cloud vendors support.
- Podman: daemonless, rootless-friendly, good systemd integration.
- nerdctl: Docker-CLI compatible, can interface containerd / k8s.
Advanced
- Multi-stage build: separate build stage from runtime, image size reduces 80%+.
- BuildKit cache:
--mount=type=cachereuses pip / apt cache, 5x faster builds. - Distroless images:
gcr.io/distroless/*images only contain runtime, minimum attack surface. - Image signing: cosign signs + verifies images, prevent supply chain attacks.