Chapter 21 β Quantization Techniques: The Full Landscape
βEvery quantization method is a different answer to the same question: how do you preserve accuracy while dramatically reducing precision?β
Overview
This chapter covers every major quantization technique for LLMs, organized by what they quantize.
Weight-Only Quantization
GPTQ (Frantar et al., 2022)
GPTQ applies Optimal Brain Quantization (OBQ) layer by layer β it quantizes each weight column while compensating the error in remaining columns using the inverse Hessian of the layerβs output.
Key ideas:
- Quantize weights one column at a time using second-order (Hessian) information
- After quantizing column j, update remaining columns to minimize total output error
- Uses a calibration dataset (~128 samples) to estimate the Hessian
Typical config: INT4, group_size=128, symmetric quantization.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
from transformers import AutoModelForCausalLM, AutoTokenizer, GPTQConfig
quantization_config = GPTQConfig(
bits=4,
group_size=128,
dataset="c4", # calibration dataset
desc_act=True, # order columns by activation magnitude
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
quantization_config=quantization_config,
device_map="auto",
)
model.save_pretrained("Llama-3.1-8B-GPTQ-INT4")
AWQ (Lin et al., 2023)
Activation-Aware Weight Quantization: not all weight channels are equally important. Channels that correspond to large activations (salient channels) should be preserved with higher precision.
Key ideas:
- Identify salient weight channels by observing activation magnitudes on calibration data
- Scale salient channels up before quantization, scale down after (equivalent to multiplying activations by inverse)
- This moves quantization difficulty from sensitive channels to insensitive ones
- No Hessian computation β faster than GPTQ
1
2
3
4
5
6
7
8
9
10
from awq import AutoAWQForCausalLM
model = AutoAWQForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B")
model.quantize(
tokenizer,
quant_config={"w_bit": 4, "q_group_size": 128, "version": "gemm"},
)
model.save_quantized("Llama-3.1-8B-AWQ")
GGUF / llama.cpp Quantization
GGUF (GPT-Generated Unified Format) is the format used by llama.cpp and Ollama. It supports many quantization variants with different quality/size tradeoffs:
| GGUF Type | Bits (effective) | Description |
|---|---|---|
| Q2_K | ~2.6 | Extreme compression, significant quality loss |
| Q3_K_M | ~3.4 | Low quality but very small |
| Q4_0 | 4.0 | Simple 4-bit, no group scaling |
| Q4_K_M | ~4.6 | 4-bit with group scales and mins. Best quality/size. |
| Q5_K_M | ~5.5 | Near-FP16 quality at ~60% size |
| Q6_K | ~6.6 | Near-lossless |
| Q8_0 | 8.0 | Essentially lossless |
1
2
3
# Convert and quantize with llama.cpp
python convert_hf_to_gguf.py meta-llama/Llama-3.1-8B --outtype f16
./llama-quantize Llama-3.1-8B-F16.gguf Llama-3.1-8B-Q4_K_M.gguf Q4_K_M
bitsandbytes
The library behind QLoRA. Two main modes:
| Mode | Method | Bits | Use Case |
|---|---|---|---|
| LLM.int8() | Mixed decomposition β FP16 for outlier channels, INT8 for rest | 8 | Inference without quality loss |
| NF4 | Normal Float 4-bit with double quantization | 4 | QLoRA fine-tuning, inference |
1
2
3
4
5
6
7
8
9
10
11
12
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
# 4-bit NF4 with double quantization (QLoRA config)
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B", quantization_config=bnb_config, device_map="auto"
)
Other Weight-Only Methods
Weight + Activation Quantization
SmoothQuant (Xiao et al., 2022)
The key insight: weights are easy to quantize; activations are hard (due to outliers). SmoothQuant migrates difficulty from activations to weights by applying per-channel scaling:
\[Y = (X \cdot \text{diag}(s)^{-1}) \cdot (\text{diag}(s) \cdot W) = \hat{X} \cdot \hat{W}\]The scaling factor $s$ balances the quantization difficulty between X and W per channel.
QServe / QoQ (W4A8KV4)
QServe (Lin et al., 2024) achieves the holy grail: 4-bit weights, 8-bit activations, AND 4-bit KV-cache in a single framework. Key technique: QoQ (Quattuor-Octo-Quattuor) progressive quantization.
| Component | Precision | Method |
|---|---|---|
| Weights | W4 | Group-quantized INT4 |
| Activations | A8 | Per-token dynamic INT8 |
| KV-Cache | KV4 | Per-channel key + per-token value INT4 |
FP8 Inference
On H100+ GPUs, FP8 inference is native and requires minimal changes:
1
2
3
4
5
6
7
8
9
10
11
12
from transformers import AutoModelForCausalLM
# FP8 inference with transformers
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
torch_dtype=torch.float8_e4m3fn, # FP8 weights
device_map="auto",
)
# Or with vLLM
from vllm import LLM
llm = LLM(model="meta-llama/Llama-3.1-8B", dtype="float16", quantization="fp8")
Quantization-Aware Training (QAT)
BitNet (Wang et al., 2023)
Trains 1-bit (ternary: -1, 0, +1) transformers from scratch. No floating-point weights at all.
- Weights: FP16 (16 bits per param)
- Matmul: FP16 Γ FP16 β FP16
- Memory: 2 bytes per parameter
- Weights: {-1, 0, +1} (1.58 bits)
- Matmul: replaces multiply with add/subtract
- Memory: ~0.2 bytes per parameter
BitNet can match FP16 quality at 3B+ params while being dramatically more efficient. The challenge: requires training from scratch.
Master Comparison Table
| Method | Target | Bits | PTQ/QAT | Calibration | Speed vs FP16 | Quality (vs FP16) | GPU Required | Best For |
|---|---|---|---|---|---|---|---|---|
| GPTQ | Weights | 4 | PTQ | ~128 samples | ~3β4Γ | -0.1β0.5 PPL | NVIDIA | GPU inference (AutoGPTQ, vLLM) |
| AWQ | Weights | 4 | PTQ | ~128 samples | ~3β4Γ | -0.1β0.3 PPL | NVIDIA | GPU inference (vLLM, TGI) |
| GGUF Q4_K_M | Weights | ~4.6 | PTQ | None | ~2β3Γ (CPU) | -0.2β0.5 PPL | CPU/GPU | llama.cpp, Ollama |
| bitsandbytes NF4 | Weights | 4 | PTQ | None | ~2Γ | -0.3β0.5 PPL | NVIDIA | QLoRA fine-tuning |
| QuIP# | Weights | 2 | PTQ | ~128 samples | ~2β3Γ | -0.5β1.0 PPL | NVIDIA | Extreme compression |
| AQLM | Weights | 2 | PTQ | Learnable | ~2Γ | -0.3β0.8 PPL | NVIDIA | 2-bit with good quality |
| HQQ | Weights | 4 | PTQ | None (zero-shot) | ~3Γ | -0.2β0.5 PPL | NVIDIA | Fast quantization |
| EXL2 | Weights | 2β8 | PTQ | ~128 samples | ~3β4Γ | Dynamic (optimal) | NVIDIA | ExLlamaV2, Pareto optimal |
| SmoothQuant | W+A | 8 | PTQ | ~512 samples | ~2Γ | -0.0β0.1 PPL | NVIDIA | W8A8 server inference |
| QServe | W+A+KV | 4/8/4 | PTQ | Calibration | ~3β4Γ | -0.1β0.3 PPL | NVIDIA | Full-stack quantization |
| FP8 | W+A | 8 | PTQ | Per-tensor | ~2Γ | Negligible | H100+ | Native H100 inference |
| BitNet | Weights | 1.58 | QAT | Full training | ~10Γ+ | Comparable at 3B+ | Any | From-scratch training |
Dequantization at Runtime
When a layer performs its computation, INT4 weights are dequantized on the fly:
Efficient implementations (GPTQ, AWQ kernels) fuse unpacking + dequantization + matmul into a single GPU kernel, avoiding materializing the full FP16 weight matrix.
Whatβs Next
With all techniques covered, the next chapter provides benchmarks and a selection guide β empirical comparisons of quality, speed, and memory across methods, plus a decision flowchart for choosing the right approach.
β Previous: Chapter 20 β Quantization Fundamentals Β· Next: Chapter 22 β Quantization Benchmarks β
Last updated: April 2026