Chapter 13 β Fine-Tuning & Adaptation
βFine-tuning is where a model goes from βknows a lotβ to βdoes what you wantβ β the difference between a base model and an assistant.β
Supervised Fine-Tuning (SFT)
SFT trains a base model on instruction-response pairs, teaching it to follow instructions and generate helpful outputs. The training objective is the same cross-entropy loss, but computed only on the response tokens (the instruction tokens are masked from the loss):
SFT Datasets
| Dataset | Size | Source | Format |
|---|---|---|---|
| OpenHermes 2.5 | 1M | GPT-4 generated | Multi-turn chat |
| UltraChat 200K | 200K | GPT-3.5/4 generated | Filtered conversations |
| Alpaca | 52K | GPT-3.5 generated | Single-turn instruction |
| FLAN Collection | 1.8M | Human + model | Diverse NLP tasks |
| SlimOrca | 517K | GPT-4 curated | Deduped multi-task |
LoRA β Low-Rank Adaptation
Full fine-tuning updates all parameters β expensive for large models (7B+ parameters). LoRA (Hu et al., 2021) decomposes weight updates into low-rank matrices, training only ~0.1β1% of parameters.
The Core Idea
Instead of updating the full weight matrix W, LoRA adds a low-rank update:
\[W' = W + \Delta W = W + \alpha \cdot B A\]where $W \in \mathbb{R}^{d \times k}$, $A \in \mathbb{R}^{r \times k}$, $B \in \mathbb{R}^{d \times r}$, and $r \ll \min(d, k)$.
- Update W directly: [d, k]
- For W_q in LLaMA-3-8B:
- 4096 Γ 4096 = 16.8M params per layer
- Memory: optimizer states for all params
- Freeze W, train A [r, k] + B [d, r]
- A: [16, 4096] = 65K params
- B: [4096, 16] = 65K params
- 130K vs 16.8M β 129Γ fewer params
LoRA Rank Selection
| Rank (r) | Trainable Params (8B model) | Quality | Use Case |
|---|---|---|---|
| 8 | ~6M | Good for simple tasks | Classification, style |
| 16 | ~13M | Good general purpose | Instruction following |
| 32 | ~26M | Better for complex tasks | Code, reasoning |
| 64 | ~52M | Near full fine-tune | Domain adaptation |
| 128 | ~104M | Diminishing returns | Rarely needed |
QLoRA β Quantized LoRA
QLoRA (Dettmers et al., 2023) combines 4-bit quantized base weights with LoRA adapters, enabling fine-tuning of large models on consumer GPUs:
Adam states: 32 GB
Gradients: 16 GB
Total: ~64 GB β needs A100
LoRA adapters (BF16): ~26 MB
Adam for LoRA: ~52 MB
Total: ~5 GB β fits on RTX 4090
Key innovations: NF4 (Normal Float 4-bit quantization, optimal for normally distributed weights), double quantization (quantize the quantization constants too), and paged optimizers (offload optimizer states to CPU).
DoRA β Weight-Decomposed LoRA
DoRA (Liu et al., 2024) decomposes weights into magnitude and direction components, applying LoRA only to the direction:
\[W' = m \cdot \frac{W + BA}{\|W + BA\|}\]This consistently outperforms LoRA at the same rank, closing the gap to full fine-tuning.
Other PEFT Methods
Full Fine-Tuning vs PEFT
| Factor | Full Fine-Tuning | LoRA/QLoRA |
|---|---|---|
| Trainable parameters | 100% | 0.1β1% |
| Memory (8B model) | 64+ GB | 5β16 GB |
| Training speed | Baseline | ~1.2β1.5Γ faster |
| Quality (general) | Best | ~95β99% of full |
| Quality (complex domain) | Best | May need higher rank |
| Multi-task | One model each | Stack/swap adapters |
| Merge into base? | N/A | Yes (W + BA) |
Fine-Tuning with TRL
TRL is the standard library for LLM fine-tuning:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import LoraConfig
from trl import SFTTrainer, SFTConfig
from datasets import load_dataset
# Load model with quantization
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B",
torch_dtype=torch.bfloat16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B")
tokenizer.pad_token = tokenizer.eos_token
# LoRA config
lora_config = LoraConfig(
r=16,
lora_alpha=32, # scaling factor: alpha/r
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
lora_dropout=0.05,
task_type="CAUSAL_LM",
)
# Dataset
dataset = load_dataset("HuggingFaceH4/ultrachat_200k", split="train_sft")
# Training
training_args = SFTConfig(
output_dir="./llama-sft",
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
learning_rate=2e-4,
num_train_epochs=1,
max_seq_length=2048,
bf16=True,
logging_steps=10,
save_strategy="steps",
save_steps=500,
)
trainer = SFTTrainer(
model=model,
args=training_args,
train_dataset=dataset,
peft_config=lora_config,
)
trainer.train()
# Merge LoRA weights back into the base model
merged_model = trainer.model.merge_and_unload()
merged_model.save_pretrained("./llama-sft-merged")
Whatβs Next
SFT trains models to follow instructions, but thereβs much more to parameter-efficient adaptation. The next chapter covers LoRA, QLoRA, DoRA, and the full landscape of PEFT methods β with detailed math and tensor shapes.
β Previous: Chapter 12 β Mid-Training & Continued Pre-Training Β· Next: Chapter 14 β PEFT: LoRA, QLoRA & Variants β
Last updated: April 2026