π§ Language Modelling β The Complete Guide
From Tokens to Frontier: Understanding, building, and deploying large language models. A comprehensive learning path for ML practitioners navigating the generative AI landscape.
Who This Is For
Youβre comfortable with Python and basic machine learning. Youβve heard about transformers, GPT, and LLMs β maybe even used them β but you want to truly understand how they work under the hood: from tokenization and attention mechanisms to quantization, CUDA kernels, and production deployment. This guide takes you from first principles all the way to the research frontier.
π Table of Contents
| # | Chapter | What Youβll Learn |
|---|---|---|
| Foundations | Β | Β |
| 1 | Introduction to Language Modelling | The full pipeline β text to tokens to tensors to predictions. Brief history, tokenization essentials, and how everything connects |
| 2 | Embeddings & Representations | From one-hot to dense vectors, positional encodings (sinusoidal, RoPE, ALiBi), and the geometry of meaning |
| 3 | The Transformer | Complete architecture walkthrough with tensor shapes at every layer β residuals, norms, feed-forward networks |
| 4 | Attention β SDPA & Multi-Head | Q/K/V math, scaled dot-product attention, multi-head mechanism, causal masking, Flash Attention |
| Attention & Architecture | Β | Β |
| 5 | Attention Variants β GQA, MQA & MLA | The KV bottleneck, grouped-query attention, multi-latent attention, sliding window β with tensor shape comparisons |
| 6 | Decoder-Only Models | GPT family, LLaMA architecture, model size tables, open-source landscape |
| 7 | Encoder & Seq2Seq Models | BERT, T5, BART β when to use encoder vs decoder vs encoder-decoder |
| 8 | Multimodal Models | Vision-language models, ViT, Whisper, modality fusion, image-to-token pipelines |
| Training | Β | Β |
| 9 | Data for LLMs & VLMs | Pre-training, mid-training, SFT, RLHF, RL data β formats, tensor shapes, and representative datasets for every phase |
| 10 | Pre-Training at Scale | Data curation, training objectives, learning rate schedules, loss curves, datasets |
| 11 | Optimizers & Loss Functions | Cross-entropy, KD loss, DPO loss; AdamW, Adam-mini, Muon, SOAP β the machinery that drives LLM training |
| 12 | Mid-Training & Continued Pre-Training | Long-context extension, domain adaptation, data mixing, annealing β the bridge between pre-training and fine-tuning |
| 13 | Fine-Tuning & Adaptation | SFT, LoRA, QLoRA, DoRA β parameter-efficient methods with rank decomposition math |
| 14 | PEFT: LoRA, QLoRA & Variants | Deep dive into LoRA math, QLoRA NF4 quantisation, DoRA, rsLoRA, VeRA, Adapters, Prefix Tuning, IAΒ³ β with tensor shapes and selection guide |
| 15 | Alignment β RLHF & Beyond | Reward models, PPO, DPO, RLAIF, Constitutional AI β making models helpful and safe |
| Inference & Generation | Β | Β |
| 16 | Inference & Sampling Strategies | Temperature, top-k/p, beam search, speculative decoding, structured output |
| 17 | KV-Cache β Mechanics & Memory | How KV-cache works, tensor shapes through generation, memory calculations, worked examples for real models |
| 18 | KV-Cache β Optimization Strategies | PagedAttention, continuous batching, eviction policies, KV-cache quantization, prefix caching |
| Numerical Precision & Quantization | Β | Β |
| 19 | Data Types & Numerical Precision | FP32 to FP4 β bit layouts, range vs precision, mixed precision training, BF16 vs FP16 |
| 20 | Quantization Fundamentals | Affine math, symmetric vs asymmetric, calibration, weight-only vs weight-activation vs KV-cache quantization |
| 21 | Quantization Techniques β Full Landscape | Every major method: GPTQ, AWQ, SmoothQuant, KIVI, BitNet + master comparison table |
| 22 | Quantization Benchmarks & Selection | Perplexity vs bits, throughput benchmarks, decision flowchart, GGUF guide |
| 23 | Knowledge Distillation & QAD | Response/feature/relation distillation, sequence-level distillation, quantisation-aware distillation, fake quantisation with STE |
| Hardware & Kernels | Β | Β |
| 24 | GPU Architecture for ML | SMs, Tensor Cores, memory hierarchy, roofline model, A100/H100/B200 comparison |
| 25 | CUDA & Kernel Development | CUDA programming model, Triton kernels, Flash Attention design, fused operations |
| 26 | Distributed Training | DDP, tensor/pipeline parallelism, FSDP, DeepSpeed ZeRO, 3D parallelism |
| 27 | ASICs & Specialized Accelerators | TPUs, Groq, Trainium, Apple Silicon β when GPUs arenβt the answer |
| Ecosystem & Tooling | Β | Β |
| 28 | The Hugging Face Ecosystem | Hub, datasets, tokenizers, Spaces β the open-source ML platform |
| 29 | Transformers Library Deep Dive | AutoModel, config, Trainer, Pipeline β internals and patterns |
| 30 | Evaluation & Benchmarks | Perplexity, MMLU, HumanEval, lm-eval-harness, Chatbot Arena |
| 31 | Serving & Deployment | vLLM, TGI, TensorRT-LLM, Ollama, llama.cpp β from prototype to production |
| Advanced Topics | Β | Β |
| 32 | Scaling Laws & Emergent Abilities | Kaplan, Chinchilla, emergence debate β the science of βhow bigβ |
| 33 | Mixture of Experts | Sparse routing, load balancing, Mixtral, DeepSeek-MoE β more params, same compute |
| 34 | SSMs & Beyond Transformers | Mamba, RWKV, Jamba β O(T) alternatives to quadratic attention |
| 35 | Reasoning Models | Chain-of-Thought, o1-style reasoning, process rewards, test-time compute scaling |
| Applications & Frontier | Β | Β |
| 36 | Retrieval-Augmented Generation | RAG pipeline, embedding models, vector databases, advanced retrieval |
| 37 | Agents & Tool Use | Function calling, coding agents, computer use, MCP protocol |
| 38 | The Frontier | Long context, world models, safety, interpretability β open problems |
| Appendices | Β | Β |
| A | History of Language Modelling | Full timeline from Shannon to GPT-4 β every milestone, paper, and breakthrough |
| B | Tokenization Deep Dive | BPE step-by-step, WordPiece, SentencePiece, training tokenizers from scratch |
| C | HuggingFace Model Config Field Reference | Every config.json parameter explained β values, formulae, caveats, VLM/VLA configs, and kernel pitfall checklist |
πΊοΈ Learning Path
β‘ Quick Start Paths
Path A: βI want to understand transformersβ (5 chapters)
- 01 β Introduction β the big picture
- 03 β The Transformer β architecture deep dive
- 04 β SDPA & Multi-Head Attention β attention mechanics
- 06 β Decoder-Only Models β GPT & LLaMA
- 16 β Inference & Sampling β how generation works
Path B: βI want to train / fine-tune modelsβ (7 chapters)
- 09 β Data for LLMs & VLMs β data across all training phases
- 10 β Pre-Training at Scale β data and objectives
- 11 β Optimizers & Loss Functions β AdamW, schedules, losses
- 12 β Mid-Training β continued pre-training
- 13 β Fine-Tuning & Adaptation β SFT & LoRA
- 14 β PEFT: LoRA, QLoRA & Variants β parameter-efficient methods
- 26 β Distributed Training β scaling to multi-GPU
Path C: βI want to deploy and optimizeβ (7 chapters)
- 17 β KV-Cache Mechanics β memory bottleneck
- 18 β KV-Cache Optimization β PagedAttention & more
- 19 β Data Types β precision tradeoffs
- 21 β Quantization Techniques β the full landscape
- 23 β Knowledge Distillation & QAD β compress with distillation
- 29 β Transformers Library β practical tooling
- 31 β Serving & Deployment β vLLM, TGI, Ollama
Path D: βI want deep understandingβ (full guide)
Read chapters 1 through 38 in order. Each builds on the previous. Appendices A and B provide additional depth on history and tokenization.
π Prerequisites
Before diving in, you should be comfortable with:
- Python β NumPy, basic PyTorch tensor operations
- Linear algebra β matrix multiplication, dot products, transposes
- ML basics β loss functions, gradient descent, backpropagation
- Probability β softmax, cross-entropy, sampling from distributions
- Command line β terminal, pip/conda, environment variables
π Changelog
| Date | Changes |
|---|---|
| April 2026 | Added Appendix C β HuggingFace Model Config Field Reference (LLM, VLM, VLA configs, kernel pitfall checklist) |
| April 2026 | Added Ch 9 (Data), Ch 11 (Optimizers & Loss Functions), Ch 14 (PEFT Deep Dive), Ch 23 (Knowledge Distillation & QAD) β renumbered all chapters accordingly |
| April 2026 | Initial release β 34 chapters + 2 appendices |
Last updated: April 2026