← Back to Table of Contents

Chapter 13 β€” Fine-Tuning & Adaptation

β€œFine-tuning is where a model goes from β€˜knows a lot’ to β€˜does what you want’ β€” the difference between a base model and an assistant.”

Supervised Fine-Tuning (SFT)

SFT trains a base model on instruction-response pairs, teaching it to follow instructions and generate helpful outputs. The training objective is the same cross-entropy loss, but computed only on the response tokens (the instruction tokens are masked from the loss):

SFT Loss Masking
System: You are a helpful assistant
User: What is attention?
Assistant: Attention is a mechanism that...
↓ Loss computation ↓
masked (no loss)
masked (no loss)
loss computed here βœ“

SFT Datasets

Dataset Size Source Format
OpenHermes 2.5 1M GPT-4 generated Multi-turn chat
UltraChat 200K 200K GPT-3.5/4 generated Filtered conversations
Alpaca 52K GPT-3.5 generated Single-turn instruction
FLAN Collection 1.8M Human + model Diverse NLP tasks
SlimOrca 517K GPT-4 curated Deduped multi-task

LoRA β€” Low-Rank Adaptation

Full fine-tuning updates all parameters β€” expensive for large models (7B+ parameters). LoRA (Hu et al., 2021) decomposes weight updates into low-rank matrices, training only ~0.1–1% of parameters.

The Core Idea

Instead of updating the full weight matrix W, LoRA adds a low-rank update:

\[W' = W + \Delta W = W + \alpha \cdot B A\]

where $W \in \mathbb{R}^{d \times k}$, $A \in \mathbb{R}^{r \times k}$, $B \in \mathbb{R}^{d \times r}$, and $r \ll \min(d, k)$.

LoRA β€” Low-Rank Decomposition
Full Fine-Tuning
  • Update W directly: [d, k]
  • For W_q in LLaMA-3-8B:
  • 4096 Γ— 4096 = 16.8M params per layer
  • Memory: optimizer states for all params
LoRA (r=16)
  • Freeze W, train A [r, k] + B [d, r]
  • A: [16, 4096] = 65K params
  • B: [4096, 16] = 65K params
  • 130K vs 16.8M β†’ 129Γ— fewer params
LoRA Tensor Shapes
Input x
[ B, T, k ]
WΒ·x (frozen)
[ B, T, d ]
original path
AΒ·x (down-project)
[ B, T, r ]
r β‰ͺ d (e.g., 16)
BΒ·(AΒ·x) (up-project)
[ B, T, d ]
LoRA delta
Output (W + BA)Β·x
[ B, T, d ]
frozen + learned combined

LoRA Rank Selection

Rank (r) Trainable Params (8B model) Quality Use Case
8 ~6M Good for simple tasks Classification, style
16 ~13M Good general purpose Instruction following
32 ~26M Better for complex tasks Code, reasoning
64 ~52M Near full fine-tune Domain adaptation
128 ~104M Diminishing returns Rarely needed

QLoRA β€” Quantized LoRA

QLoRA (Dettmers et al., 2023) combines 4-bit quantized base weights with LoRA adapters, enabling fine-tuning of large models on consumer GPUs:

QLoRA Memory Savings
Full Fine-Tuning (LLaMA-3-8B)
BF16 weights: 16 GB
Adam states: 32 GB
Gradients: 16 GB
Total: ~64 GB β†’ needs A100
QLoRA (LLaMA-3-8B)
NF4 weights: 4 GB
LoRA adapters (BF16): ~26 MB
Adam for LoRA: ~52 MB
Total: ~5 GB β†’ fits on RTX 4090

Key innovations: NF4 (Normal Float 4-bit quantization, optimal for normally distributed weights), double quantization (quantize the quantization constants too), and paged optimizers (offload optimizer states to CPU).

DoRA β€” Weight-Decomposed LoRA

DoRA (Liu et al., 2024) decomposes weights into magnitude and direction components, applying LoRA only to the direction:

\[W' = m \cdot \frac{W + BA}{\|W + BA\|}\]

This consistently outperforms LoRA at the same rank, closing the gap to full fine-tuning.

Other PEFT Methods

Parameter-Efficient Fine-Tuning Landscape
Adapters
Small bottleneck layers inserted after attention/FFN. ~2–5% trainable params. Adds latency.
Prefix Tuning
Prepend learnable "soft prompts" to keys/values at each layer. No weight changes to model.
Prompt Tuning
Prepend learnable embeddings to input. Simplest PEFT β€” only ~100K params. Works at scale.
IAΒ³
Learn scaling vectors for K, V, and FFN. Even fewer params than LoRA. Good for few-shot.

Full Fine-Tuning vs PEFT

Factor Full Fine-Tuning LoRA/QLoRA
Trainable parameters 100% 0.1–1%
Memory (8B model) 64+ GB 5–16 GB
Training speed Baseline ~1.2–1.5Γ— faster
Quality (general) Best ~95–99% of full
Quality (complex domain) Best May need higher rank
Multi-task One model each Stack/swap adapters
Merge into base? N/A Yes (W + BA)

Fine-Tuning with TRL

TRL is the standard library for LLM fine-tuning:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import LoraConfig
from trl import SFTTrainer, SFTConfig
from datasets import load_dataset

# Load model with quantization
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B")
tokenizer.pad_token = tokenizer.eos_token

# LoRA config
lora_config = LoraConfig(
    r=16,
    lora_alpha=32,               # scaling factor: alpha/r
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
    lora_dropout=0.05,
    task_type="CAUSAL_LM",
)

# Dataset
dataset = load_dataset("HuggingFaceH4/ultrachat_200k", split="train_sft")

# Training
training_args = SFTConfig(
    output_dir="./llama-sft",
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    learning_rate=2e-4,
    num_train_epochs=1,
    max_seq_length=2048,
    bf16=True,
    logging_steps=10,
    save_strategy="steps",
    save_steps=500,
)

trainer = SFTTrainer(
    model=model,
    args=training_args,
    train_dataset=dataset,
    peft_config=lora_config,
)
trainer.train()

# Merge LoRA weights back into the base model
merged_model = trainer.model.merge_and_unload()
merged_model.save_pretrained("./llama-sft-merged")

What’s Next

SFT trains models to follow instructions, but there’s much more to parameter-efficient adaptation. The next chapter covers LoRA, QLoRA, DoRA, and the full landscape of PEFT methods β€” with detailed math and tensor shapes.

← Previous: Chapter 12 β€” Mid-Training & Continued Pre-Training Β· Next: Chapter 14 β€” PEFT: LoRA, QLoRA & Variants β†’


Last updated: April 2026