Skip to main content
ZICQ

Wiki Concepts

CUDA

Concepts
Aliases: CUDA GPU programming NVIDIA ·2026-09-14

CUDA

CUDA is NVIDIA's GPU parallel computing API. Deep learning depends almost entirely on CUDA (AMD ROCm and Apple Metal are catching up but ecosystem is an order of magnitude behind).

Why GPUs suit deep learning

  • High parallelism: an A100 has 6912 CUDA cores, runs thousands of threads simultaneously.
  • Matrix multiplication blazing fast: Tensor Core computes 16x16x16 FP16 matmul per beat.
  • High memory bandwidth: HBM3 reaches 2 TB/s on A100, 20x faster than CPU DDR5.
  • Mature ecosystem: PyTorch / TensorFlow / JAX all natively support CUDA.

Mainstream GPU specs

Model Memory FP16 TFLOPS Use
H100 80 GB 989 Training / inference flagship
A100 40/80 GB 312 Mainstream training
L40S 48 GB 366 Inference dedicated
RTX 4090 24 GB 165 Consumer / small models
RTX 5090 32 GB 209 New consumer flagship
B200 192 GB 2250 Next-gen flagship

Software stack

  • CUDA Toolkit: compiler (nvcc), runtime, math libraries (cuBLAS, cuDNN).
  • cuDNN: high-performance implementation of deep learning primitives (conv, attention, layer norm).
  • NCCL: multi-GPU / multi-node communication (AllReduce).
  • TensorRT: inference optimizer (kernel fusion, quantization, graph optimization).
  • Triton: NVIDIA's Python-like GPU programming language, PyTorch 2.0 integrated.

Why LLM training needs GPUs

Training a 70B model takes ~1500 H100s for 1 month. Just forward + backward needs dozens of matrix multiplies per token. CPU completely infeasible.

Why LLM inference needs GPUs too

  • Matrix multiplication: each decode step needs attention / FFN matmul per token.
  • Large memory: 70B FP16 model = 140 GB, KV cache takes another tens of GB.
  • Small batch latency sensitive: single-user streaming, GPU low-latency advantage irreplaceable.

AMD / Apple / Intel attempts

  • AMD ROCm: CUDA open-source alternative, HIP translation layer. But ecosystem lagging, hardware/software issues.
  • Apple Metal / MLX: Apple Silicon unified memory advantage (no CPU↔GPU copy), decent performance but compute ceiling low (M3 Max ~26 TFLOPS, 40x weaker than H100).
  • Intel Gaudi / oneAPI: training can match A100, ecosystem weaker.

Practical advice

  • NVIDIA still first choice: most mature ecosystem, vLLM / Triton / Megatron-LM all default NVIDIA.
  • Align CUDA versions: PyTorch / vLLM each require specific CUDA versions, mixing is tricky.
  • VRAM not enough fallbacks: CPU offload, quantization (INT4/INT8), ZeRO-3, DeepSpeed.
  • Monitor GPU utilization: nvidia-smi dmon or DCGM, avoid GPU idle waste.