← Back to Table of Contents

Chapter 20 β€” Quantization Fundamentals

β€œQuantization is the art of throwing away precision strategically β€” keeping what matters and discarding what doesn’t.”

What Is Quantization?

Quantization maps high-precision floating-point values to lower-precision integers. For LLMs, this typically means FP16 β†’ INT8 or INT4, reducing memory and enabling faster inference.

\[x_q = \text{round}\!\left(\frac{x}{s}\right) + z\]

where $s$ = scale factor, $z$ = zero point, and $x_q$ is the quantized integer value.

Dequantization (at inference time): $\hat{x} = s \cdot (x_q - z)$

Symmetric vs Asymmetric Quantization

Symmetric versus asymmetric INT8 quantization showing number lines, formulas, and worked examples
Symmetric quantization maps to [βˆ’127, 127] around zero; asymmetric maps to [0, 255] with a zero-point offset
Symmetric vs Asymmetric Quantization
Symmetric
  • Zero point z = 0 (no offset)
  • Scale: s = max(|x|) / (2^(b-1) - 1)
  • Range: [-127, 127] for INT8
  • Simpler, faster dequantization
  • Wastes range if distribution is skewed
Asymmetric
  • Zero point z β‰  0 (shifted)
  • Scale: s = (max(x) - min(x)) / (2^b - 1)
  • Zero point: z = round(-min(x) / s)
  • Full range utilization
  • Better for skewed distributions
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
import torch

def symmetric_quantize(x, bits=8):
    """Symmetric quantization β€” zero-centered."""
    qmax = 2 ** (bits - 1) - 1                    # 127 for INT8
    scale = x.abs().max() / qmax
    x_q = torch.round(x / scale).clamp(-qmax, qmax).to(torch.int8)
    return x_q, scale

def symmetric_dequantize(x_q, scale):
    """Dequantize back to float."""
    return x_q.float() * scale

# Example
w = torch.randn(4096, 4096)                       # FP32 weight matrix
w_q, scale = symmetric_quantize(w, bits=8)
w_hat = symmetric_dequantize(w_q, scale)           # Reconstructed
print(f"Original: {w.nbytes / 1e6:.1f} MB")       # 67.1 MB
print(f"Quantized: {w_q.nbytes / 1e6:.1f} MB")    # 16.8 MB (4Γ— smaller)
print(f"Max error: {(w - w_hat).abs().max():.6f}") # Small reconstruction error

Quantization Granularity

Where you compute scales dramatically affects quality:

Quantization Granularity Levels
Per-Tensor: 1 scale for entire weight matrix fastest, lowest quality
Per-Channel: 1 scale per output channel (row) good balance
Per-Group: 1 scale per group of G values (e.g., G=128) best quality for INT4
Per-Element: 1 scale per value no compression β€” pointless
Per-Group Quantization Tensor Layout
Original weight
[ d_out, d_in ] @ FP16: 2 bytes each
Quantized weight
[ d_out, d_in/pack ] @ INT4: 8 values packed per int32
Scales
[ d_out, d_in/G ] @ FP16: 1 scale per group of G
Zero points
[ d_out, d_in/G ] @ INT4 or FP16

Effective bits per value with group size 128: INT4 + scales + zeros β‰ˆ 4.25 bits (the overhead from storing scales is small).

PTQ vs QAT

Post-Training Quantization vs Quantization-Aware Training
PTQ (Post-Training)
  • Quantize after training is complete
  • Uses calibration dataset (~128–512 samples)
  • No additional training needed
  • Fast: minutes to hours
  • Lower quality at extreme compression (INT3/2)
QAT (Aware Training)
  • Simulate quantization during training
  • Model adapts to quantization noise
  • Requires training infrastructure
  • Slow: full training cost
  • Best quality at any bit-width

For LLMs, PTQ dominates because models are expensive to retrain. QAT methods like BitNet train from scratch.

Calibration Strategies

PTQ methods need a calibration dataset to determine quantization parameters (scales, zero points). The choice of calibration strategy matters:

Strategy Description Quality
Min-Max Use observed min/max of activations Baseline
Percentile Clip to 99.9th percentile (ignore outliers) Better
MSE Minimize reconstruction error (MSE) per layer Best
Entropy Minimize KL divergence between original and quantized Good

What Gets Quantized

Quantization Targets
Weight-Only
Quantize weights to INT4/8. Activations stay in FP16. Simplest, most common. Memory-bound speedup.
Weight + Activation
Quantize both weights and activations (W8A8, W4A8). Enables INT8 matmul. Compute-bound speedup.
KV-Cache
Quantize cached K, V tensors. Reduces memory for long sequences. See Ch 15. Serving-bound savings.

When to use each:

  • Weight-only: most LLM inference (decode is memory-bandwidth bound β€” smaller weights = faster reads)
  • Weight + activation: latency-sensitive serving, batch inference (prefill is compute-bound)
  • KV-cache: long-context serving, high-throughput scenarios

The Outlier Problem

LLM activations have massive outliers β€” a few channels with values 100Γ— larger than average. These outliers make naive quantization fail because the scale is dominated by outliers, leaving most values underrepresented.

1
2
3
4
5
6
7
# Typical activation distribution in a transformer
# 99% of values: [-3, 3]
# 0.1% outlier channels: [-100, 100]
# If we quantize to [-127, 127] with max=100:
#   scale = 100/127 β‰ˆ 0.79
#   A value of 1.0 maps to round(1.0/0.79) = 1 β†’ dequant = 0.79
#   Quantization error of 0.21 β€” over 20% for typical values!

This is why methods like SmoothQuant, GPTQ, and AWQ were invented β€” they’re all different strategies for handling outliers. See Chapter 21.

What’s Next

With the fundamentals in place, the next chapter surveys the full landscape of quantization techniques β€” from GPTQ and AWQ to SmoothQuant and BitNet, with a master comparison table.

← Previous: Chapter 19 β€” Data Types & Numerical Precision Β· Next: Chapter 21 β€” Quantization Techniques β†’


Last updated: April 2026