Skip to main content
ZICQ

Wiki Concepts

Model Quantization

Concepts
Aliases: GPTQ AWQ GGUF ·2026-09-14

Model Quantization

Quantization converts model weights from high precision (FP32/FP16) to lower precision (INT8/INT4/FP8), dramatically reducing memory footprint and inference latency. The cost is slight accuracy loss.

Precision comparison

| Precision | Bytes per param | | 7B model size | |---|---|---| | FP32 | 4 | 28 GB | | FP16/BF16 | 2 | 14 GB | | INT8 | 1 | 7 GB | | INT4 | 0.5 | 3.5 GB |

Mainstream approaches

  • GPTQ: INT4, GPU inference first choice. text-generation-inference / AutoGPTQ.
  • AWQ (Activation-aware Weight Quantization): INT4, preserves precision on activation-sensitive weights, often better than GPTQ on small models.
  • GGUF (GPT-Generated Unified Format): CPU / Apple Silicon / mixed devices, native to llama.cpp.
  • BNB (BitsAndBytes): INT8/INT4, easiest HuggingFace integration, used by QLoRA.
  • FP8: native on H100/newer cards, near-zero accuracy loss.

Practical recommendations

  • Tight memory: choose INT4 (GPTQ / AWQ / GGUF Q4).
  • Accuracy first: choose INT8 (GGUF Q8).
  • CPU/Apple Silicon: choose GGUF.
  • NVIDIA 4090/5090: AWQ or BNB INT4.