Chapter 7 β Encoder & Encoder-Decoder Models
βWhile decoder-only models dominate generation, encoders remain the workhorses of NLU β classification, retrieval, and structured extraction are still best served by bidirectional context.β
Encoder-Only: BERT
BERT (Devlin et al., 2018) processes the entire input bidirectionally β every token can attend to every other token. No causal mask, no generation β just rich contextual representations.
Masked Language Modelling (MLM)
BERTβs pre-training objective: randomly mask 15% of tokens and predict them from context.
BERT Variants
| Model | Params | Layers | Hidden | Heads | Max Length |
|---|---|---|---|---|---|
| BERT-base | 110M | 12 | 768 | 12 | 512 |
| BERT-large | 340M | 24 | 1024 | 16 | 512 |
| RoBERTa | 355M | 24 | 1024 | 16 | 512 |
| DeBERTa-v3 | 304M | 24 | 1024 | 16 | 512 |
| ModernBERT | 395M | 28 | 1024 | 16 | 8192 |
RoBERTa improved on BERT by removing the NSP objective, training longer, and using larger batches. DeBERTa added disentangled attention (separate content and position). ModernBERT (2024) brought modern architectural improvements (RoPE, GeLU, Flash Attention) and longer context.
BERT for Classification
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
tokenizer = AutoTokenizer.from_pretrained("microsoft/deberta-v3-base")
model = AutoModelForSequenceClassification.from_pretrained(
"microsoft/deberta-v3-base",
num_labels=2, # binary classification
)
inputs = tokenizer("This movie was fantastic!", return_tensors="pt", padding=True)
# inputs["input_ids"].shape: [1, T]
outputs = model(**inputs)
# outputs.logits.shape: [1, 2] β one score per class
prediction = torch.argmax(outputs.logits, dim=-1)
print(f"Predicted class: {prediction.item()}") # 0 or 1
Encoder-Decoder: T5
T5 (Raffel et al., 2019) frames every NLP task as text-to-text. The encoder processes the input, and the decoder generates the output autoregressively, using cross-attention to read from encoder representations.
- Input: "translate English to French: Hello world"
- Full bidirectional self-attention
- Output: contextual representations [B, T_enc, d]
- Output: "Bonjour le monde"
- Causal self-attention + cross-attention to encoder
- Cross-attn: Q from decoder, K/V from encoder
The crucial difference from decoder-only: the cross-attention layer lets the decoder attend to arbitrary encoder positions (not just preceding tokens), giving it direct access to the full input representation.
T5 for Summarization
1
2
3
4
5
6
7
8
9
10
11
12
13
14
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("google-t5/t5-base")
model = AutoModelForSeq2SeqLM.from_pretrained("google-t5/t5-base")
text = """summarize: The transformer architecture has revolutionized natural
language processing. Originally proposed for machine translation, it has since
been adapted for virtually every NLP task, from text classification to
open-ended generation."""
inputs = tokenizer(text, return_tensors="pt", max_length=512, truncation=True)
outputs = model.generate(inputs.input_ids, max_new_tokens=64)
summary = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(summary)
BART: Denoising Autoencoder
BART (Lewis et al., 2019) is another encoder-decoder model, pre-trained by corrupting input text (token masking, deletion, permutation, infilling) and learning to reconstruct the original. This makes it particularly strong for summarization and text infilling.
| Corruption | Description |
|---|---|
| Token Masking | Replace tokens with [MASK] |
| Token Deletion | Randomly delete tokens |
| Text Infilling | Replace spans with single [MASK] |
| Sentence Permutation | Shuffle sentence order |
When to Use Which Architecture
Models: BERT, RoBERTa, DeBERTa
Key: Bidirectional context, [CLS] pooling
Models: T5, BART, mBART
Key: Cross-attention bridges input and output
Models: GPT, LLaMA, Mistral
Key: Scales best, dominant paradigm
| Task | Best Architecture | Why |
|---|---|---|
| Text classification | Encoder | Bidirectional context captures full meaning |
| Named Entity Recognition | Encoder | Per-token classification needs bidirectional |
| Semantic similarity | Encoder | Sentence embeddings from [CLS] or mean pooling |
| Translation | Encoder-Decoder | Input and output are different languages/structures |
| Summarization | Encoder-Decoder or Decoder | Both work; decoder-only now competitive |
| Open-ended chat | Decoder-only | Autoregressive generation is natural |
| Code generation | Decoder-only | Long-range dependencies, large pre-training |
| Reasoning / math | Decoder-only | Chain-of-thought requires sequential generation |
Trend: Decoder-only models are increasingly used for tasks that were previously encoder territory, especially at larger scales. But for embedding, classification, and retrieval at smaller scales, encoders remain more efficient and often more accurate.
Whatβs Next
Modern LLMs donβt just process text β they handle images, audio, and video. The next chapter explores multimodal models and how different modalities are projected into a shared representation space.
β Previous: Chapter 6 β Decoder-Only Models Β· Next: Chapter 8 β Multimodal Models β
Last updated: April 2026