Chapter 8 β Multimodal Models
βA language model that can see, hear, and reason across modalities isnβt just a better chatbot β itβs a step toward general-purpose intelligence.β
The Multimodal Idea
Modern LLMs are no longer text-only. Vision-Language Models (VLMs) process images alongside text, audio models handle speech, and emerging architectures handle video, 3D, and more. The key challenge is projecting each modality into the shared representation space of the language model.
Vision Encoders
Vision Transformers (ViT) convert images into sequences of patch embeddings, making them directly compatible with transformer architectures.
How ViT Works
[3, 336, 336]
14Γ14 β 576 patches
[576, d_vision]
[576, d_vision]
[576, d_vision]
Common Vision Encoders
| Encoder | Training | Resolution | Patches | d_vision | Used By |
|---|---|---|---|---|---|
| CLIP ViT-L/14 | Contrastive (image-text pairs) | 224 | 256 | 1024 | LLaVA 1.0 |
| SigLIP SO400M | Sigmoid contrastive | 384 | 729 | 1152 | LLaVA-NeXT, PaliGemma |
| InternViT-6B | Contrastive + generative | 448 | 1024 | 3200 | InternVL 2 |
| DINOv2 ViT-L | Self-supervised | 518 | 1369 | 1024 | Various |
Fusion Strategies
How vision and language tokens interact defines the model architecture:
Key VLM Architectures
LLaVA (Visual Instruction Tuning)
The most influential open-source VLM architecture β elegantly simple.
Training: (1) Pre-train projector on image-caption pairs (freeze vision encoder + LLM), then (2) instruction-tune the LLM + projector on visual QA data.
Other Architectures
| Model | Vision Encoder | Projection | LLM | Key Innovation |
|---|---|---|---|---|
| LLaVA-NeXT | SigLIP | 2-layer MLP | LLaMA/Vicuna | AnyRes (dynamic resolution) |
| GPT-4V/4o | Proprietary | Proprietary | GPT-4 | Seamless multimodal reasoning |
| Gemini | Natively multimodal | Shared encoder | Gemini | Trained multimodal from scratch |
| PaliGemma | SigLIP | Linear | Gemma | Small, efficient, compositional |
| Qwen2-VL | ViT + M-RoPE | MLP | Qwen2 | Multimodal RoPE, dynamic resolution |
| Pixtral | Custom 400M ViT | β | Mistral | Variable resolution, no padding |
Audio Models: Whisper
Whisper (OpenAI, 2022) is an encoder-decoder model for speech recognition:
| Whisper Model | Params | Layers (Enc/Dec) | d_model | Heads |
|---|---|---|---|---|
| tiny | 39M | 4/4 | 384 | 6 |
| base | 74M | 6/6 | 512 | 8 |
| small | 244M | 12/12 | 768 | 12 |
| medium | 769M | 24/24 | 1024 | 16 |
| large-v3 | 1.5B | 32/32 | 1280 | 20 |
Using a VLM with Transformers
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
from transformers import AutoProcessor, AutoModelForVision2Seq
from PIL import Image
import torch
model_id = "llava-hf/llava-v1.6-mistral-7b-hf"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForVision2Seq.from_pretrained(
model_id, torch_dtype=torch.float16, device_map="auto"
)
image = Image.open("example.jpg")
prompt = "<image>\nDescribe this image in detail."
inputs = processor(text=prompt, images=image, return_tensors="pt").to(model.device)
# inputs["pixel_values"].shape: [1, 3, 336, 336]
# inputs["input_ids"].shape: [1, T_text] (includes <image> placeholder)
output = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(output[0], skip_special_tokens=True))
The Token Budget Problem
Each image becomes hundreds of tokens (576 for a 336Γ336 image with 14Γ14 patches). With multiple images or higher resolution, image tokens can dominate the context window:
| Resolution | Patch Size | Tokens per Image |
|---|---|---|
| 224Γ224 | 14Γ14 | 256 |
| 336Γ336 | 14Γ14 | 576 |
| 448Γ448 | 14Γ14 | 1024 |
| Dynamic (4 tiles) | 14Γ14 | ~2304 |
Solutions: token compression (average pooling, learned resampling), dynamic resolution (process at native aspect ratio, tile into sub-images), and early fusion with pooling (Qwen2-VL reduces tokens with a perceiver-style resampler).
Whatβs Next
With architectures covered β text-only (encoder, decoder, encoder-decoder) and multimodal β we now shift to how these models are trained. The next chapter covers pre-training at scale.
β Previous: Chapter 7 β Encoder & Seq2Seq Models Β· Next: Chapter 9 β Data for LLMs β
Last updated: April 2026