LoRA (Low-Rank Adaptation)
LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning (PEFT) method: freeze the base model, insert low-rank matrices as trainable "side paths" in each layer. Original paper: LoRA: Low-Rank Adaptation of Large Language Models (2021).
Core idea
Weight updates ΔW don't need full rank. Low-rank decomposition ΔW = A·B (A is d×r, B is r×k, r ≪ d,k) suffices. Train only A and B, base weights stay frozen.
Why it works
- Save memory: full-parameter FT on 7B needs ~60GB VRAM; LoRA needs only ~16GB.
- Pluggable: a base model can host multiple LoRA adapters, switched on demand.
- No inference overhead: after training, LoRA weights can be merged back into the base model — zero inference cost.
Key hyperparameters
- rank (r): rank of the low-rank matrices. Common: 8/16/32/64, larger = stronger but slower.
- alpha: LoRA scaling factor, usually set to 2×rank.
- target_modules: which layers to attach LoRA to (typically q_proj / v_proj, or all).
Advanced
- QLoRA: 4-bit quantized base + LoRA, fits 70B on a single 24G GPU.
- AdaLoRA: adaptively allocates rank per layer.
- DoRA: weight-decomposed variant, more stable than LoRA.