← Back to Table of Contents

Chapter 7 β€” Encoder & Encoder-Decoder Models

β€œWhile decoder-only models dominate generation, encoders remain the workhorses of NLU β€” classification, retrieval, and structured extraction are still best served by bidirectional context.”

Encoder-Only: BERT

BERT (Devlin et al., 2018) processes the entire input bidirectionally β€” every token can attend to every other token. No causal mask, no generation β€” just rich contextual representations.

Masked Language Modelling (MLM)

BERT’s pre-training objective: randomly mask 15% of tokens and predict them from context.

BERT β€” Masked Language Modelling
The
[MASK]
sat
on
the
[MASK]
↓ Bidirectional Transformer Encoder ↓
β€”
cat
β€”
β€”
β€”
mat
BERT Tensor Flow
Input
[ B, T ]
token IDs
Embeddings
[ B, T, 768 ]
token + position + segment
Encoder output
[ B, T, 768 ]
contextual representations
[CLS] pooled
[ B, 768 ]
sentence-level representation

BERT Variants

Model Params Layers Hidden Heads Max Length
BERT-base 110M 12 768 12 512
BERT-large 340M 24 1024 16 512
RoBERTa 355M 24 1024 16 512
DeBERTa-v3 304M 24 1024 16 512
ModernBERT 395M 28 1024 16 8192

RoBERTa improved on BERT by removing the NSP objective, training longer, and using larger batches. DeBERTa added disentangled attention (separate content and position). ModernBERT (2024) brought modern architectural improvements (RoPE, GeLU, Flash Attention) and longer context.

BERT for Classification

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch

tokenizer = AutoTokenizer.from_pretrained("microsoft/deberta-v3-base")
model = AutoModelForSequenceClassification.from_pretrained(
    "microsoft/deberta-v3-base",
    num_labels=2,  # binary classification
)

inputs = tokenizer("This movie was fantastic!", return_tensors="pt", padding=True)
# inputs["input_ids"].shape: [1, T]

outputs = model(**inputs)
# outputs.logits.shape: [1, 2]  ← one score per class

prediction = torch.argmax(outputs.logits, dim=-1)
print(f"Predicted class: {prediction.item()}")  # 0 or 1

Encoder-Decoder: T5

T5 (Raffel et al., 2019) frames every NLP task as text-to-text. The encoder processes the input, and the decoder generates the output autoregressively, using cross-attention to read from encoder representations.

T5 β€” Text-to-Text Framework
Encoder (Bidirectional)
  • Input: "translate English to French: Hello world"
  • Full bidirectional self-attention
  • Output: contextual representations [B, T_enc, d]
Decoder (Causal)
  • Output: "Bonjour le monde"
  • Causal self-attention + cross-attention to encoder
  • Cross-attn: Q from decoder, K/V from encoder
Encoder-Decoder Tensor Flow
Encoder Input: [B, T_enc] β†’ Encoder β†’ [B, T_enc, d_model]
Decoder Input: [B, T_dec] β†’ Embed β†’ [B, T_dec, d_model]
Decoder Layer: Causal Self-Attention [B, T_dec, d] β†’ Cross-Attention(Q=dec, K/V=enc) β†’ FFN
Output Logits: [B, T_dec, vocab_size]

The crucial difference from decoder-only: the cross-attention layer lets the decoder attend to arbitrary encoder positions (not just preceding tokens), giving it direct access to the full input representation.

T5 for Summarization

1
2
3
4
5
6
7
8
9
10
11
12
13
14
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("google-t5/t5-base")
model = AutoModelForSeq2SeqLM.from_pretrained("google-t5/t5-base")

text = """summarize: The transformer architecture has revolutionized natural
language processing. Originally proposed for machine translation, it has since
been adapted for virtually every NLP task, from text classification to
open-ended generation."""

inputs = tokenizer(text, return_tensors="pt", max_length=512, truncation=True)
outputs = model.generate(inputs.input_ids, max_new_tokens=64)
summary = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(summary)

BART: Denoising Autoencoder

BART (Lewis et al., 2019) is another encoder-decoder model, pre-trained by corrupting input text (token masking, deletion, permutation, infilling) and learning to reconstruct the original. This makes it particularly strong for summarization and text infilling.

Corruption Description
Token Masking Replace tokens with [MASK]
Token Deletion Randomly delete tokens
Text Infilling Replace spans with single [MASK]
Sentence Permutation Shuffle sentence order

When to Use Which Architecture

Architecture Selection Guide
Encoder-Only
Best for: Classification, NER, retrieval, semantic similarity
Models: BERT, RoBERTa, DeBERTa
Key: Bidirectional context, [CLS] pooling
Encoder-Decoder
Best for: Translation, summarization, structured generation
Models: T5, BART, mBART
Key: Cross-attention bridges input and output
Decoder-Only
Best for: Open-ended generation, chat, code, reasoning
Models: GPT, LLaMA, Mistral
Key: Scales best, dominant paradigm
Task Best Architecture Why
Text classification Encoder Bidirectional context captures full meaning
Named Entity Recognition Encoder Per-token classification needs bidirectional
Semantic similarity Encoder Sentence embeddings from [CLS] or mean pooling
Translation Encoder-Decoder Input and output are different languages/structures
Summarization Encoder-Decoder or Decoder Both work; decoder-only now competitive
Open-ended chat Decoder-only Autoregressive generation is natural
Code generation Decoder-only Long-range dependencies, large pre-training
Reasoning / math Decoder-only Chain-of-thought requires sequential generation

Trend: Decoder-only models are increasingly used for tasks that were previously encoder territory, especially at larger scales. But for embedding, classification, and retrieval at smaller scales, encoders remain more efficient and often more accurate.

What’s Next

Modern LLMs don’t just process text β€” they handle images, audio, and video. The next chapter explores multimodal models and how different modalities are projected into a shared representation space.

← Previous: Chapter 6 β€” Decoder-Only Models Β· Next: Chapter 8 β€” Multimodal Models β†’


Last updated: April 2026