Chapter 31 β Serving & Deployment
βTraining a great model is half the battle. Serving it efficiently β low latency, high throughput, at reasonable cost β is the other half.β
Inference Frameworks Landscape
vLLM
The most popular open-source LLM serving engine. Key innovations from Chapter 18:
- PagedAttention: Non-contiguous KV-cache with virtual memory β no fragmentation
- Continuous batching: New requests join mid-batch instead of waiting
- Prefix caching: Reuse KV-cache for shared prefixes (system prompts)
- Speculative decoding: Use a small draft model for faster generation
- Quantization: GPTQ, AWQ, FP8, bitsandbytes support
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
# vLLM β Python API
from vllm import LLM, SamplingParams
llm = LLM(
model="meta-llama/Llama-3.1-8B-Instruct",
dtype="bfloat16",
tensor_parallel_size=2, # split across 2 GPUs
max_model_len=8192,
gpu_memory_utilization=0.9,
enable_prefix_caching=True,
)
params = SamplingParams(temperature=0.7, top_p=0.9, max_tokens=512)
outputs = llm.generate(["Explain transformers in one paragraph."], params)
print(outputs[0].outputs[0].text)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
# vLLM β OpenAI-compatible server
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--dtype bfloat16 \
--tensor-parallel-size 2 \
--max-model-len 8192 \
--port 8000
# Use with any OpenAI SDK client
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [{"role": "user", "content": "Hello!"}],
"temperature": 0.7,
"max_tokens": 256
}'
Text Generation Inference (TGI)
HuggingFaceβs production server, used by the Inference API:
1
2
3
4
5
6
7
8
9
10
11
12
13
# TGI via Docker
docker run --gpus all --shm-size 1g -p 8080:80 \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id meta-llama/Llama-3.1-8B-Instruct \
--quantize awq \
--max-input-length 4096 \
--max-total-tokens 8192 \
--max-batch-prefill-tokens 4096
# Query
curl http://localhost:8080/generate \
-H "Content-Type: application/json" \
-d '{"inputs": "What is attention?", "parameters": {"max_new_tokens": 200}}'
TensorRT-LLM
NVIDIAβs highest-performance option, using custom CUDA kernels:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
# Build TensorRT-LLM engine
# Step 1: Convert HF model to TensorRT-LLM checkpoint
python convert_checkpoint.py \
--model_dir meta-llama/Llama-3.1-8B-Instruct \
--output_dir ./trt_ckpt \
--dtype bfloat16 \
--tp_size 2
# Step 2: Build engine with optimizations
trtllm-build \
--checkpoint_dir ./trt_ckpt \
--output_dir ./trt_engine \
--gemm_plugin bfloat16 \
--max_batch_size 64 \
--max_input_len 4096 \
--max_seq_len 8192 \
--use_fp8_context_fmha enable
# Step 3: Serve via Triton Inference Server
llama.cpp & Ollama
For local deployment on CPU and Apple Silicon:
1
2
3
4
5
6
7
8
9
# Ollama β simplest way to run models locally
ollama pull llama3.1:8b-instruct-q4_K_M
ollama run llama3.1:8b-instruct-q4_K_M
# Python client
import ollama
response = ollama.chat(model='llama3.1:8b-instruct-q4_K_M', messages=[
{'role': 'user', 'content': 'Explain attention in one sentence.'}
])
1
2
3
4
5
6
# llama.cpp β direct usage with GGUF files
./llama-cli \
-m models/llama-3.1-8b-instruct-Q4_K_M.gguf \
-p "Explain attention:" \
-n 256 \
-ngl 99 # offload all layers to GPU (Metal on macOS)
Framework Comparison
| Feature | vLLM | TGI | TensorRT-LLM | llama.cpp | SGLang |
|---|---|---|---|---|---|
| PagedAttention | β | β | β (custom) | β | β |
| Continuous batching | β | β | β | β | β |
| Tensor parallelism | β | β | β | β | β |
| Quantization | GPTQ, AWQ, FP8, bnb | GPTQ, AWQ, EETQ | FP8, INT8, INT4 | GGUF (2-8 bit) | GPTQ, AWQ, FP8 |
| Speculative decode | β | β | β | β | β |
| Structured output | Via outlines | β (grammar) | β | β (grammar) | β (native) |
| OpenAI API | β | β | Via Triton | β | β |
| CPU inference | β | β | β | β | β |
| Apple Silicon | β | β | β | β (Metal) | β |
| Best for | General GPU serving | HF ecosystem prod | Maximum NVIDIA perf | Local / edge | Structured output |
Latency vs Throughput
These are often conflicting goals:
- Small batch size (1β4)
- Tensor parallelism across GPUs
- Speculative decoding
- FP8 or INT4 quantization
- Use case: chatbot, real-time
- Large batch size (64β256)
- Continuous batching
- Prefix caching for shared prompts
- Maximize GPU utilization
- Use case: batch processing, APIs
Key metrics:
- Time to First Token (TTFT): How long until the first token appears (prefill latency)
- Time per Output Token (TPOT): Average decode latency per token
- Tokens per Second (TPS): Total throughput across all concurrent requests
Deployment Checklist
- Choose quantization: W4A16 (AWQ/GPTQ) if quality-sensitive, FP8 if hardware supports it, GGUF for CPU/Apple
- Choose framework: vLLM for most GPU deployments, Ollama for local
- Set max_model_len: Donβt set higher than needed β it reserves KV-cache memory
- Enable prefix caching: If system prompt is shared across requests
- Monitor: Track TTFT, TPOT, queue depth, GPU utilization, KV-cache usage
- Load test: Use benchmarking tools to find the throughput/latency sweet spot
Whatβs Next
From infrastructure, we now move to the science of scale. The next chapter explores scaling laws β the mathematical relationships between compute, data, parameters, and performance.
β Previous: Chapter 30 β Evaluation & Benchmarks Β· Next: Chapter 32 β Scaling Laws & Emergent Abilities β
Last updated: April 2026