← Back to Table of Contents

Chapter 21 β€” Quantization Techniques: The Full Landscape

β€œEvery quantization method is a different answer to the same question: how do you preserve accuracy while dramatically reducing precision?”

Overview

This chapter covers every major quantization technique for LLMs, organized by what they quantize.

Quantization Technique Map
Weight-Only Quantization
GPTQ, AWQ, GGUF, bitsandbytes, QuIP#, AQLM, HQQ, EXL2
Weight + Activation Quantization
SmoothQuant, QServe/QoQ, Atom, SpinQuant, FP8
QAT (Train with Quantization)
LLM-QAT, BitNet, OneBit

Weight-Only Quantization

GPTQ (Frantar et al., 2022)

GPTQ applies Optimal Brain Quantization (OBQ) layer by layer β€” it quantizes each weight column while compensating the error in remaining columns using the inverse Hessian of the layer’s output.

Key ideas:

  • Quantize weights one column at a time using second-order (Hessian) information
  • After quantizing column j, update remaining columns to minimize total output error
  • Uses a calibration dataset (~128 samples) to estimate the Hessian

Typical config: INT4, group_size=128, symmetric quantization.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
from transformers import AutoModelForCausalLM, AutoTokenizer, GPTQConfig

quantization_config = GPTQConfig(
    bits=4,
    group_size=128,
    dataset="c4",            # calibration dataset
    desc_act=True,           # order columns by activation magnitude
)

model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B",
    quantization_config=quantization_config,
    device_map="auto",
)
model.save_pretrained("Llama-3.1-8B-GPTQ-INT4")

AWQ (Lin et al., 2023)

Activation-Aware Weight Quantization: not all weight channels are equally important. Channels that correspond to large activations (salient channels) should be preserved with higher precision.

Key ideas:

  • Identify salient weight channels by observing activation magnitudes on calibration data
  • Scale salient channels up before quantization, scale down after (equivalent to multiplying activations by inverse)
  • This moves quantization difficulty from sensitive channels to insensitive ones
  • No Hessian computation β€” faster than GPTQ
1
2
3
4
5
6
7
8
9
10
from awq import AutoAWQForCausalLM

model = AutoAWQForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B")

model.quantize(
    tokenizer,
    quant_config={"w_bit": 4, "q_group_size": 128, "version": "gemm"},
)
model.save_quantized("Llama-3.1-8B-AWQ")

GGUF / llama.cpp Quantization

GGUF (GPT-Generated Unified Format) is the format used by llama.cpp and Ollama. It supports many quantization variants with different quality/size tradeoffs:

GGUF Type Bits (effective) Description
Q2_K ~2.6 Extreme compression, significant quality loss
Q3_K_M ~3.4 Low quality but very small
Q4_0 4.0 Simple 4-bit, no group scaling
Q4_K_M ~4.6 4-bit with group scales and mins. Best quality/size.
Q5_K_M ~5.5 Near-FP16 quality at ~60% size
Q6_K ~6.6 Near-lossless
Q8_0 8.0 Essentially lossless
1
2
3
# Convert and quantize with llama.cpp
python convert_hf_to_gguf.py meta-llama/Llama-3.1-8B --outtype f16
./llama-quantize Llama-3.1-8B-F16.gguf Llama-3.1-8B-Q4_K_M.gguf Q4_K_M

bitsandbytes

The library behind QLoRA. Two main modes:

Mode Method Bits Use Case
LLM.int8() Mixed decomposition β€” FP16 for outlier channels, INT8 for rest 8 Inference without quality loss
NF4 Normal Float 4-bit with double quantization 4 QLoRA fine-tuning, inference
1
2
3
4
5
6
7
8
9
10
11
12
from transformers import AutoModelForCausalLM, BitsAndBytesConfig

# 4-bit NF4 with double quantization (QLoRA config)
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B", quantization_config=bnb_config, device_map="auto"
)

Other Weight-Only Methods

QuIP# (Tseng et al., 2024)
Incoherence processing + lattice codebook quantization. State-of-the-art at 2-bit. Randomized Hadamard transforms make weights "incoherent" (uniform magnitude).
AQLM (Egiazarian et al., 2024)
Additive quantization with learned codebooks. Multiple codebooks are summed to approximate each weight group. Best at extreme compression (2-bit).
HQQ (Badri & Shaji, 2023)
Half-Quadratic Quantization. Zero-shot β€” no calibration data needed. Very fast quantization. Good quality at 4-bit.
EXL2 (turboderp, 2023)
ExLlamaV2 format. Dynamic per-layer bit allocation β€” important layers get more bits. Optimal Pareto frontier for quality vs size.

Weight + Activation Quantization

SmoothQuant (Xiao et al., 2022)

The key insight: weights are easy to quantize; activations are hard (due to outliers). SmoothQuant migrates difficulty from activations to weights by applying per-channel scaling:

\[Y = (X \cdot \text{diag}(s)^{-1}) \cdot (\text{diag}(s) \cdot W) = \hat{X} \cdot \hat{W}\]

The scaling factor $s$ balances the quantization difficulty between X and W per channel.

SmoothQuant β€” Migrating Outlier Difficulty
Original: X has outliers, W is smooth activations hard to quantize
Per-channel scaling: divide X by s, multiply W by s mathematically equivalent
Smoothed: XΜ‚ is smooth, Ε΄ slightly harder both are now quantizable
W8A8: INT8 weights Γ— INT8 activations β†’ INT8 matmul 2Γ— speedup over FP16

QServe / QoQ (W4A8KV4)

QServe (Lin et al., 2024) achieves the holy grail: 4-bit weights, 8-bit activations, AND 4-bit KV-cache in a single framework. Key technique: QoQ (Quattuor-Octo-Quattuor) progressive quantization.

Component Precision Method
Weights W4 Group-quantized INT4
Activations A8 Per-token dynamic INT8
KV-Cache KV4 Per-channel key + per-token value INT4

FP8 Inference

On H100+ GPUs, FP8 inference is native and requires minimal changes:

1
2
3
4
5
6
7
8
9
10
11
12
from transformers import AutoModelForCausalLM

# FP8 inference with transformers
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B",
    torch_dtype=torch.float8_e4m3fn,  # FP8 weights
    device_map="auto",
)

# Or with vLLM
from vllm import LLM
llm = LLM(model="meta-llama/Llama-3.1-8B", dtype="float16", quantization="fp8")

Quantization-Aware Training (QAT)

BitNet (Wang et al., 2023)

Trains 1-bit (ternary: -1, 0, +1) transformers from scratch. No floating-point weights at all.

BitNet β€” Ternary Weight Transformer
Standard Transformer
  • Weights: FP16 (16 bits per param)
  • Matmul: FP16 Γ— FP16 β†’ FP16
  • Memory: 2 bytes per parameter
BitNet b1.58
  • Weights: {-1, 0, +1} (1.58 bits)
  • Matmul: replaces multiply with add/subtract
  • Memory: ~0.2 bytes per parameter

BitNet can match FP16 quality at 3B+ params while being dramatically more efficient. The challenge: requires training from scratch.


Master Comparison Table

Method Target Bits PTQ/QAT Calibration Speed vs FP16 Quality (vs FP16) GPU Required Best For
GPTQ Weights 4 PTQ ~128 samples ~3–4Γ— -0.1–0.5 PPL NVIDIA GPU inference (AutoGPTQ, vLLM)
AWQ Weights 4 PTQ ~128 samples ~3–4Γ— -0.1–0.3 PPL NVIDIA GPU inference (vLLM, TGI)
GGUF Q4_K_M Weights ~4.6 PTQ None ~2–3Γ— (CPU) -0.2–0.5 PPL CPU/GPU llama.cpp, Ollama
bitsandbytes NF4 Weights 4 PTQ None ~2Γ— -0.3–0.5 PPL NVIDIA QLoRA fine-tuning
QuIP# Weights 2 PTQ ~128 samples ~2–3Γ— -0.5–1.0 PPL NVIDIA Extreme compression
AQLM Weights 2 PTQ Learnable ~2Γ— -0.3–0.8 PPL NVIDIA 2-bit with good quality
HQQ Weights 4 PTQ None (zero-shot) ~3Γ— -0.2–0.5 PPL NVIDIA Fast quantization
EXL2 Weights 2–8 PTQ ~128 samples ~3–4Γ— Dynamic (optimal) NVIDIA ExLlamaV2, Pareto optimal
SmoothQuant W+A 8 PTQ ~512 samples ~2Γ— -0.0–0.1 PPL NVIDIA W8A8 server inference
QServe W+A+KV 4/8/4 PTQ Calibration ~3–4Γ— -0.1–0.3 PPL NVIDIA Full-stack quantization
FP8 W+A 8 PTQ Per-tensor ~2Γ— Negligible H100+ Native H100 inference
BitNet Weights 1.58 QAT Full training ~10Γ—+ Comparable at 3B+ Any From-scratch training

Dequantization at Runtime

When a layer performs its computation, INT4 weights are dequantized on the fly:

INT4 Dequantization Flow
Packed INT4 weights: [d_out, d_in/8] 8 values per int32
Unpack: extract 4-bit values β†’ [d_out, d_in] INT4 integers
Dequantize: (int4 - zero) Γ— scale β†’ FP16 [d_out, d_in] @ FP16
Matmul: FP16 weights Γ— FP16 activations β†’ FP16 output

Efficient implementations (GPTQ, AWQ kernels) fuse unpacking + dequantization + matmul into a single GPU kernel, avoiding materializing the full FP16 weight matrix.

What’s Next

With all techniques covered, the next chapter provides benchmarks and a selection guide β€” empirical comparisons of quality, speed, and memory across methods, plus a decision flowchart for choosing the right approach.

← Previous: Chapter 20 β€” Quantization Fundamentals Β· Next: Chapter 22 β€” Quantization Benchmarks β†’


Last updated: April 2026