Chapter 30 β Evaluation & Benchmarks
βIf you canβt measure it, you canβt improve it. But if you measure the wrong thing, you optimize for the wrong outcome.β
Perplexity
The most fundamental language model metric β how surprised is the model by the test data?
\[\text{PPL} = \exp\left(-\frac{1}{N}\sum_{i=1}^{N} \log p(x_i | x_{<i})\right)\]- Lower is better β a perfect model has PPL = 1
- Typical values: GPT-2 ~29 on WikiText-103, LLaMA-3-8B ~6
- Only meaningful for comparing models on the same tokenizer + dataset
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
def compute_perplexity(model, tokenizer, text, max_length=2048, stride=512):
"""Compute perplexity with sliding window."""
encodings = tokenizer(text, return_tensors="pt")
input_ids = encodings.input_ids.to(model.device)
seq_len = input_ids.size(1)
nlls = []
for begin in range(0, seq_len, stride):
end = min(begin + max_length, seq_len)
target_len = end - (begin if begin == 0 else begin + max_length - stride)
input_chunk = input_ids[:, begin:end]
with torch.no_grad():
outputs = model(input_chunk, labels=input_chunk)
nlls.append(outputs.loss * target_len)
if end == seq_len:
break
ppl = torch.exp(torch.stack(nlls).sum() / seq_len)
return ppl.item()
Major Benchmarks
LLM Benchmark Categories
Knowledge & Reasoning
MMLU, ARC, HellaSwag, WinoGrande, TruthfulQA, GPQA
Math & Code
GSM8K, MATH, HumanEval, MBPP, LiveCodeBench
Chat & Instruction
MT-Bench, AlpacaEval, Chatbot Arena (ELO), WildBench
Knowledge & Reasoning Benchmarks
| Benchmark | Tasks | Metric | What It Tests |
|---|---|---|---|
| MMLU | 57 subjects, 14K questions | Accuracy | Broad academic knowledge (STEM, humanities, social sciences) |
| MMLU-Pro | Harder MMLU with 10 choices | Accuracy | Deeper reasoning with reduced guessing |
| ARC (Challenge) | Grade-school science, 2.5K | Accuracy | Scientific reasoning |
| HellaSwag | 10K sentence completions | Accuracy | Common-sense reasoning, adversarial |
| WinoGrande | 44K pronoun resolution | Accuracy | Coreference / common sense |
| TruthfulQA | 817 questions | MC accuracy | Resistance to common misconceptions |
| GPQA | Graduate-level science | Accuracy | Expert knowledge (PhD difficulty) |
Math & Code Benchmarks
| Benchmark | Format | Metric | What It Tests |
|---|---|---|---|
| GSM8K | 8.5K grade-school math word problems | Solve rate | Multi-step arithmetic reasoning |
| MATH | 12.5K competition math problems (5 levels) | Solve rate | Advanced mathematical reasoning |
| HumanEval | 164 Python function completions | pass@k | Code generation correctness |
| MBPP | 974 Python problems | pass@k | Basic Python programming |
| LiveCodeBench | Continuously updated contest problems | pass@k | Non-contaminated coding ability |
Chat & Human Evaluation
| Benchmark | Method | Metric | What It Tests |
|---|---|---|---|
| MT-Bench | GPT-4 judges 80 multi-turn conversations | 1β10 score | Instruction following, multi-turn coherence |
| AlpacaEval 2 | GPT-4 compares outputs to reference | Win rate (LC) | Instruction following quality |
| Chatbot Arena | Humans vote on blind pairwise comparisons | ELO rating | Overall chat preference (gold standard) |
| WildBench | Real user queries, LLM-judged | Win rate | Challenging real-world tasks |
The Open LLM Leaderboard
HuggingFaceβs Open LLM Leaderboard evaluates models using lm-eval-harness:
Leaderboard v2 suite (as of 2024):
- MMLU-Pro (knowledge)
- GPQA (expert knowledge)
- MuSR (multi-step reasoning)
- MATH (mathematical reasoning)
- BBH (Big-Bench Hard β diverse reasoning)
- IFEval (instruction following)
Running Evaluations with lm-eval-harness
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
# Install
pip install lm-eval
# Evaluate a model on MMLU
lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-3.1-8B-Instruct \
--tasks mmlu \
--batch_size auto \
--num_fewshot 5 \
--output_path results/
# Multiple benchmarks
lm_eval --model hf \
--model_args pretrained=meta-llama/Llama-3.1-8B-Instruct,dtype=bfloat16 \
--tasks mmlu,hellaswag,arc_challenge,winogrande,gsm8k \
--batch_size auto \
--output_path results/
# Evaluate a GPTQ quantized model
lm_eval --model hf \
--model_args pretrained=TheBloke/Llama-3-8B-GPTQ,autogptq=True \
--tasks mmlu \
--batch_size auto
Benchmark Limitations
Known Issues with Benchmarks
Contamination
- Training data may contain benchmark questions
- Models memorize answers rather than reason
- Especially problematic for MMLU, HellaSwag
- LiveCodeBench addresses this with fresh problems
Goodhart's Law
- "When a measure becomes a target, it ceases to be a good measure"
- Models are optimized for benchmarks, not real tasks
- High MMLU β good chatbot
- Chatbot Arena (human pref) remains the most trusted signal
Other limitations:
- Sensitivity to prompting: Few-shot count, prompt format, and chat template all affect scores significantly
- Saturation: Some benchmarks become too easy β HellaSwag is >95% for most modern models
- Narrow scope: Benchmarks donβt test creativity, nuance, or long-form coherence
- Static datasets: Real-world capability evolves, benchmarks donβt
Practical Evaluation Strategy
For your own fine-tuned models:
- Perplexity on a held-out set β did training reduce loss?
- Task-specific metrics β accuracy, F1, exact match on your use case
- lm-eval-harness on standard benchmarks β to compare with baselines
- Qualitative inspection β read 50 outputs manually. There is no substitute
- Chatbot Arena (if applicable) β compare against known models
Whatβs Next
Evaluating models is one thing β serving them at scale is another. The next chapter covers serving and deployment with vLLM, TGI, and other inference frameworks.
β Previous: Chapter 29 β Transformers Library Deep Dive Β· Next: Chapter 31 β Serving & Deployment β
Last updated: April 2026