CUDA
CUDA is NVIDIA's GPU parallel computing API. Deep learning depends almost entirely on CUDA (AMD ROCm and Apple Metal are catching up but ecosystem is an order of magnitude behind).
Why GPUs suit deep learning
- High parallelism: an A100 has 6912 CUDA cores, runs thousands of threads simultaneously.
- Matrix multiplication blazing fast: Tensor Core computes 16x16x16 FP16 matmul per beat.
- High memory bandwidth: HBM3 reaches 2 TB/s on A100, 20x faster than CPU DDR5.
- Mature ecosystem: PyTorch / TensorFlow / JAX all natively support CUDA.
Mainstream GPU specs
| Model | Memory | FP16 TFLOPS | Use |
|---|---|---|---|
| H100 | 80 GB | 989 | Training / inference flagship |
| A100 | 40/80 GB | 312 | Mainstream training |
| L40S | 48 GB | 366 | Inference dedicated |
| RTX 4090 | 24 GB | 165 | Consumer / small models |
| RTX 5090 | 32 GB | 209 | New consumer flagship |
| B200 | 192 GB | 2250 | Next-gen flagship |
Software stack
- CUDA Toolkit: compiler (nvcc), runtime, math libraries (cuBLAS, cuDNN).
- cuDNN: high-performance implementation of deep learning primitives (conv, attention, layer norm).
- NCCL: multi-GPU / multi-node communication (AllReduce).
- TensorRT: inference optimizer (kernel fusion, quantization, graph optimization).
- Triton: NVIDIA's Python-like GPU programming language, PyTorch 2.0 integrated.
Why LLM training needs GPUs
Training a 70B model takes ~1500 H100s for 1 month. Just forward + backward needs dozens of matrix multiplies per token. CPU completely infeasible.
Why LLM inference needs GPUs too
- Matrix multiplication: each decode step needs attention / FFN matmul per token.
- Large memory: 70B FP16 model = 140 GB, KV cache takes another tens of GB.
- Small batch latency sensitive: single-user streaming, GPU low-latency advantage irreplaceable.
AMD / Apple / Intel attempts
- AMD ROCm: CUDA open-source alternative, HIP translation layer. But ecosystem lagging, hardware/software issues.
- Apple Metal / MLX: Apple Silicon unified memory advantage (no CPU↔GPU copy), decent performance but compute ceiling low (M3 Max ~26 TFLOPS, 40x weaker than H100).
- Intel Gaudi / oneAPI: training can match A100, ecosystem weaker.
Practical advice
- NVIDIA still first choice: most mature ecosystem, vLLM / Triton / Megatron-LM all default NVIDIA.
- Align CUDA versions: PyTorch / vLLM each require specific CUDA versions, mixing is tricky.
- VRAM not enough fallbacks: CPU offload, quantization (INT4/INT8), ZeRO-3, DeepSpeed.
- Monitor GPU utilization:
nvidia-smi dmonor DCGM, avoid GPU idle waste.