Model Quantization
Quantization converts model weights from high precision (FP32/FP16) to lower precision (INT8/INT4/FP8), dramatically reducing memory footprint and inference latency. The cost is slight accuracy loss.
Precision comparison
| Precision | Bytes per param | | 7B model size | |---|---|---| | FP32 | 4 | 28 GB | | FP16/BF16 | 2 | 14 GB | | INT8 | 1 | 7 GB | | INT4 | 0.5 | 3.5 GB |
Mainstream approaches
- GPTQ: INT4, GPU inference first choice.
text-generation-inference/AutoGPTQ. - AWQ (Activation-aware Weight Quantization): INT4, preserves precision on activation-sensitive weights, often better than GPTQ on small models.
- GGUF (GPT-Generated Unified Format): CPU / Apple Silicon / mixed devices, native to
llama.cpp. - BNB (BitsAndBytes): INT8/INT4, easiest HuggingFace integration, used by QLoRA.
- FP8: native on H100/newer cards, near-zero accuracy loss.
Practical recommendations
- Tight memory: choose INT4 (GPTQ / AWQ / GGUF Q4).
- Accuracy first: choose INT8 (GGUF Q8).
- CPU/Apple Silicon: choose GGUF.
- NVIDIA 4090/5090: AWQ or BNB INT4.