← Back to Table of Contents

Chapter 8 β€” Multimodal Models

β€œA language model that can see, hear, and reason across modalities isn’t just a better chatbot β€” it’s a step toward general-purpose intelligence.”

The Multimodal Idea

Modern LLMs are no longer text-only. Vision-Language Models (VLMs) process images alongside text, audio models handle speech, and emerging architectures handle video, 3D, and more. The key challenge is projecting each modality into the shared representation space of the language model.

Multimodal LLM β€” General Architecture
Image: [B, C, H, W] e.g., [1, 3, 336, 336]
Vision Encoder (ViT/SigLIP): [B, N_patches, d_vision] e.g., [1, 576, 1024]
Projection Layer: [B, N_patches, d_model] map vision β†’ LLM space
Concatenate: [B, N_patches + T_text, d_model] vision tokens + text tokens
LLM Decoder: process combined sequence β†’ generate text

Vision Encoders

Vision Transformers (ViT) convert images into sequences of patch embeddings, making them directly compatible with transformer architectures.

How ViT Works

Vision Transformer β€” Image to Patch Tokens
Image
[3, 336, 336]
Split into patches
14Γ—14 β†’ 576 patches
Linear embed
[576, d_vision]
+ position embed
[576, d_vision]
ViT Encoder
[576, d_vision]
Vision Encoder Tensor Shapes
Raw image
[ B, 3, 336, 336 ]
RGB pixels
Patches (14Γ—14 px)
[ B, 576, 588 ]
14Γ—14Γ—3 = 588 per patch
After ViT encoder
[ B, 576, 1024 ]
d_vision (SigLIP-L)
After projection
[ B, 576, 4096 ]
mapped to LLM d_model

Common Vision Encoders

Encoder Training Resolution Patches d_vision Used By
CLIP ViT-L/14 Contrastive (image-text pairs) 224 256 1024 LLaVA 1.0
SigLIP SO400M Sigmoid contrastive 384 729 1152 LLaVA-NeXT, PaliGemma
InternViT-6B Contrastive + generative 448 1024 3200 InternVL 2
DINOv2 ViT-L Self-supervised 518 1369 1024 Various

Fusion Strategies

How vision and language tokens interact defines the model architecture:

Modality Fusion Approaches
Early Fusion
Concatenate vision + text tokens before feeding to LLM. Vision tokens attend to text and vice versa in every layer. Most common.
Cross-Attention Fusion
Add cross-attention layers where text queries attend to vision keys/values. Keeps modalities partially separate. Used by Flamingo.
Late Fusion
Process modalities independently, combine only at decision time. Simplest but weakest interaction.

Key VLM Architectures

LLaVA (Visual Instruction Tuning)

The most influential open-source VLM architecture β€” elegantly simple.

LLaVA Architecture
Image β†’ Vision Encoder (CLIP/SigLIP) β†’ [B, N, d_vision]
MLP Projector: [B, N, d_vision] β†’ [B, N, d_model] 2-layer MLP
Interleave: [system tokens] [image tokens] [user text tokens]
LLM (Vicuna / LLaMA): process combined sequence β†’ generate response

Training: (1) Pre-train projector on image-caption pairs (freeze vision encoder + LLM), then (2) instruction-tune the LLM + projector on visual QA data.

Other Architectures

Model Vision Encoder Projection LLM Key Innovation
LLaVA-NeXT SigLIP 2-layer MLP LLaMA/Vicuna AnyRes (dynamic resolution)
GPT-4V/4o Proprietary Proprietary GPT-4 Seamless multimodal reasoning
Gemini Natively multimodal Shared encoder Gemini Trained multimodal from scratch
PaliGemma SigLIP Linear Gemma Small, efficient, compositional
Qwen2-VL ViT + M-RoPE MLP Qwen2 Multimodal RoPE, dynamic resolution
Pixtral Custom 400M ViT β€” Mistral Variable resolution, no padding

Audio Models: Whisper

Whisper (OpenAI, 2022) is an encoder-decoder model for speech recognition:

Whisper Architecture
Audio β†’ Mel Spectrogram: [B, 80, 3000] 80 mel bins, 30s max
2Γ— Conv1D β†’ [B, 1500, d_model] downsample 2Γ—
Transformer Encoder (bidirectional) β†’ [B, 1500, d_model]
Transformer Decoder (causal + cross-attention) β†’ text tokens
Whisper Model Params Layers (Enc/Dec) d_model Heads
tiny 39M 4/4 384 6
base 74M 6/6 512 8
small 244M 12/12 768 12
medium 769M 24/24 1024 16
large-v3 1.5B 32/32 1280 20

Using a VLM with Transformers

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
from transformers import AutoProcessor, AutoModelForVision2Seq
from PIL import Image
import torch

model_id = "llava-hf/llava-v1.6-mistral-7b-hf"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForVision2Seq.from_pretrained(
    model_id, torch_dtype=torch.float16, device_map="auto"
)

image = Image.open("example.jpg")
prompt = "<image>\nDescribe this image in detail."

inputs = processor(text=prompt, images=image, return_tensors="pt").to(model.device)
# inputs["pixel_values"].shape: [1, 3, 336, 336]
# inputs["input_ids"].shape:    [1, T_text]  (includes <image> placeholder)

output = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(output[0], skip_special_tokens=True))

The Token Budget Problem

Each image becomes hundreds of tokens (576 for a 336Γ—336 image with 14Γ—14 patches). With multiple images or higher resolution, image tokens can dominate the context window:

Resolution Patch Size Tokens per Image
224Γ—224 14Γ—14 256
336Γ—336 14Γ—14 576
448Γ—448 14Γ—14 1024
Dynamic (4 tiles) 14Γ—14 ~2304

Solutions: token compression (average pooling, learned resampling), dynamic resolution (process at native aspect ratio, tile into sub-images), and early fusion with pooling (Qwen2-VL reduces tokens with a perceiver-style resampler).

What’s Next

With architectures covered β€” text-only (encoder, decoder, encoder-decoder) and multimodal β€” we now shift to how these models are trained. The next chapter covers pre-training at scale.

← Previous: Chapter 7 β€” Encoder & Seq2Seq Models Β· Next: Chapter 9 β€” Data for LLMs β†’


Last updated: April 2026