Skip to main content
ZICQ

Wiki Infrastructure

Docker

Infrastructure
Aliases: Docker container containerization ·2026-09-14

Docker

Docker is the de facto standard containerization platform. Packages applications

  • dependencies + config into images, build once, run anywhere. Mainstream form for LLM service deployment (vllm / ollama / microservices).

Why LLM projects all use Docker

  • Environment consistency: same image across dev / test / prod, avoid "works on my machine".
  • GPU passthrough: --gpus all directly uses host NVIDIA driver, no driver installation in container.
  • Model weights mount: large model weights via volume mount, images stay slim.
  • Fast replication: spin up multiple dev / staging / prod instances with one command.

Core concepts

  • Image: read-only template (layered structure, shared base layers).
  • Container: running instance of an image.
  • Dockerfile: build script for an image.
  • Volume: persistent data / config files mount.
  • Network: container-to-container network (bridge / host / overlay).
  • Registry: image repository (Docker Hub, GHCR, private Harbor / ECR).

Typical Dockerfile for LLM service

FROM nvcr.io/nvidia/pytorch:24.05-py3  # Pre-installed CUDA + PyTorch
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
EXPOSE 8000
CMD ["python", "-m", "vllm.entrypoints.openai.api_server", \
     "--model", "/models/llama-3.1-70b", \
     "--tensor-parallel-size", "2"]

docker-compose multi-service orchestration

services:
  vllm:
    image: vllm/vllm-openai:latest
    runtime: nvidia  # or deploy.resources.reservations.devices
    volumes:
      - /data/models:/models
    ports:
      - "8000:8000"
  
  app:
    build: .
    depends_on: [vllm, postgres]
    environment:
      VLLM_BASE_URL: http://vllm:8000

Practical experience

  • Don't install NVIDIA driver in container: use --gpus passthrough, CUDA toolkit in image, driver on host.
  • Multi-stage build for smaller image: build stage installs dev deps, runtime stage only COPYs build output.
  • Run as non-root user: USER 10001, avoid container escape attacks.
  • Read-only rootfs: container root partition read-only, config / data via volume, minimize attack surface.
  • Health checks: each service configure HEALTHCHECK, compose / k8s determines readiness from it.

Docker vs Podman / nerdctl

  • Docker: most mature ecosystem, all cloud vendors support.
  • Podman: daemonless, rootless-friendly, good systemd integration.
  • nerdctl: Docker-CLI compatible, can interface containerd / k8s.

Advanced

  • Multi-stage build: separate build stage from runtime, image size reduces 80%+.
  • BuildKit cache: --mount=type=cache reuses pip / apt cache, 5x faster builds.
  • Distroless images: gcr.io/distroless/* images only contain runtime, minimum attack surface.
  • Image signing: cosign signs + verifies images, prevent supply chain attacks.