🧠 Language Modelling β€” The Complete Guide

From Tokens to Frontier: Understanding, building, and deploying large language models. A comprehensive learning path for ML practitioners navigating the generative AI landscape.

Who This Is For

You’re comfortable with Python and basic machine learning. You’ve heard about transformers, GPT, and LLMs β€” maybe even used them β€” but you want to truly understand how they work under the hood: from tokenization and attention mechanisms to quantization, CUDA kernels, and production deployment. This guide takes you from first principles all the way to the research frontier.

πŸ“‹ Table of Contents

# Chapter What You’ll Learn
Foundations Β  Β 
1 Introduction to Language Modelling The full pipeline β€” text to tokens to tensors to predictions. Brief history, tokenization essentials, and how everything connects
2 Embeddings & Representations From one-hot to dense vectors, positional encodings (sinusoidal, RoPE, ALiBi), and the geometry of meaning
3 The Transformer Complete architecture walkthrough with tensor shapes at every layer β€” residuals, norms, feed-forward networks
4 Attention β€” SDPA & Multi-Head Q/K/V math, scaled dot-product attention, multi-head mechanism, causal masking, Flash Attention
Attention & Architecture Β  Β 
5 Attention Variants β€” GQA, MQA & MLA The KV bottleneck, grouped-query attention, multi-latent attention, sliding window β€” with tensor shape comparisons
6 Decoder-Only Models GPT family, LLaMA architecture, model size tables, open-source landscape
7 Encoder & Seq2Seq Models BERT, T5, BART β€” when to use encoder vs decoder vs encoder-decoder
8 Multimodal Models Vision-language models, ViT, Whisper, modality fusion, image-to-token pipelines
Training Β  Β 
9 Data for LLMs & VLMs Pre-training, mid-training, SFT, RLHF, RL data β€” formats, tensor shapes, and representative datasets for every phase
10 Pre-Training at Scale Data curation, training objectives, learning rate schedules, loss curves, datasets
11 Optimizers & Loss Functions Cross-entropy, KD loss, DPO loss; AdamW, Adam-mini, Muon, SOAP β€” the machinery that drives LLM training
12 Mid-Training & Continued Pre-Training Long-context extension, domain adaptation, data mixing, annealing β€” the bridge between pre-training and fine-tuning
13 Fine-Tuning & Adaptation SFT, LoRA, QLoRA, DoRA β€” parameter-efficient methods with rank decomposition math
14 PEFT: LoRA, QLoRA & Variants Deep dive into LoRA math, QLoRA NF4 quantisation, DoRA, rsLoRA, VeRA, Adapters, Prefix Tuning, IAΒ³ β€” with tensor shapes and selection guide
15 Alignment β€” RLHF & Beyond Reward models, PPO, DPO, RLAIF, Constitutional AI β€” making models helpful and safe
Inference & Generation Β  Β 
16 Inference & Sampling Strategies Temperature, top-k/p, beam search, speculative decoding, structured output
17 KV-Cache β€” Mechanics & Memory How KV-cache works, tensor shapes through generation, memory calculations, worked examples for real models
18 KV-Cache β€” Optimization Strategies PagedAttention, continuous batching, eviction policies, KV-cache quantization, prefix caching
Numerical Precision & Quantization Β  Β 
19 Data Types & Numerical Precision FP32 to FP4 β€” bit layouts, range vs precision, mixed precision training, BF16 vs FP16
20 Quantization Fundamentals Affine math, symmetric vs asymmetric, calibration, weight-only vs weight-activation vs KV-cache quantization
21 Quantization Techniques β€” Full Landscape Every major method: GPTQ, AWQ, SmoothQuant, KIVI, BitNet + master comparison table
22 Quantization Benchmarks & Selection Perplexity vs bits, throughput benchmarks, decision flowchart, GGUF guide
23 Knowledge Distillation & QAD Response/feature/relation distillation, sequence-level distillation, quantisation-aware distillation, fake quantisation with STE
Hardware & Kernels Β  Β 
24 GPU Architecture for ML SMs, Tensor Cores, memory hierarchy, roofline model, A100/H100/B200 comparison
25 CUDA & Kernel Development CUDA programming model, Triton kernels, Flash Attention design, fused operations
26 Distributed Training DDP, tensor/pipeline parallelism, FSDP, DeepSpeed ZeRO, 3D parallelism
27 ASICs & Specialized Accelerators TPUs, Groq, Trainium, Apple Silicon β€” when GPUs aren’t the answer
Ecosystem & Tooling Β  Β 
28 The Hugging Face Ecosystem Hub, datasets, tokenizers, Spaces β€” the open-source ML platform
29 Transformers Library Deep Dive AutoModel, config, Trainer, Pipeline β€” internals and patterns
30 Evaluation & Benchmarks Perplexity, MMLU, HumanEval, lm-eval-harness, Chatbot Arena
31 Serving & Deployment vLLM, TGI, TensorRT-LLM, Ollama, llama.cpp β€” from prototype to production
Advanced Topics Β  Β 
32 Scaling Laws & Emergent Abilities Kaplan, Chinchilla, emergence debate β€” the science of β€œhow big”
33 Mixture of Experts Sparse routing, load balancing, Mixtral, DeepSeek-MoE β€” more params, same compute
34 SSMs & Beyond Transformers Mamba, RWKV, Jamba β€” O(T) alternatives to quadratic attention
35 Reasoning Models Chain-of-Thought, o1-style reasoning, process rewards, test-time compute scaling
Applications & Frontier Β  Β 
36 Retrieval-Augmented Generation RAG pipeline, embedding models, vector databases, advanced retrieval
37 Agents & Tool Use Function calling, coding agents, computer use, MCP protocol
38 The Frontier Long context, world models, safety, interpretability β€” open problems
Appendices Β  Β 
A History of Language Modelling Full timeline from Shannon to GPT-4 β€” every milestone, paper, and breakthrough
B Tokenization Deep Dive BPE step-by-step, WordPiece, SentencePiece, training tokenizers from scratch
C HuggingFace Model Config Field Reference Every config.json parameter explained β€” values, formulae, caveats, VLM/VLA configs, and kernel pitfall checklist

πŸ—ΊοΈ Learning Path

Recommended Learning Path
πŸ“– Ch 1–4: Foundations & Transformer
🧩 Ch 5–8: Attention Variants & Architectures
πŸ‹οΈ Ch 9–15: Training Pipeline
⚑ Ch 16–18: Inference & KV-Cache
πŸ”’ Ch 19–23: Precision, Quantization & Distillation
πŸ–₯️ Ch 24–27: Hardware & Kernels
πŸ› οΈ Ch 28–31: Ecosystem & Deployment
πŸš€ Ch 32–38: Advanced & Frontier

⚑ Quick Start Paths

Path A: β€œI want to understand transformers” (5 chapters)

  1. 01 β€” Introduction β€” the big picture
  2. 03 β€” The Transformer β€” architecture deep dive
  3. 04 β€” SDPA & Multi-Head Attention β€” attention mechanics
  4. 06 β€” Decoder-Only Models β€” GPT & LLaMA
  5. 16 β€” Inference & Sampling β€” how generation works

Path B: β€œI want to train / fine-tune models” (7 chapters)

  1. 09 β€” Data for LLMs & VLMs β€” data across all training phases
  2. 10 β€” Pre-Training at Scale β€” data and objectives
  3. 11 β€” Optimizers & Loss Functions β€” AdamW, schedules, losses
  4. 12 β€” Mid-Training β€” continued pre-training
  5. 13 β€” Fine-Tuning & Adaptation β€” SFT & LoRA
  6. 14 β€” PEFT: LoRA, QLoRA & Variants β€” parameter-efficient methods
  7. 26 β€” Distributed Training β€” scaling to multi-GPU

Path C: β€œI want to deploy and optimize” (7 chapters)

  1. 17 β€” KV-Cache Mechanics β€” memory bottleneck
  2. 18 β€” KV-Cache Optimization β€” PagedAttention & more
  3. 19 β€” Data Types β€” precision tradeoffs
  4. 21 β€” Quantization Techniques β€” the full landscape
  5. 23 β€” Knowledge Distillation & QAD β€” compress with distillation
  6. 29 β€” Transformers Library β€” practical tooling
  7. 31 β€” Serving & Deployment β€” vLLM, TGI, Ollama

Path D: β€œI want deep understanding” (full guide)

Read chapters 1 through 38 in order. Each builds on the previous. Appendices A and B provide additional depth on history and tokenization.

πŸ“š Prerequisites

Before diving in, you should be comfortable with:

  • Python β€” NumPy, basic PyTorch tensor operations
  • Linear algebra β€” matrix multiplication, dot products, transposes
  • ML basics β€” loss functions, gradient descent, backpropagation
  • Probability β€” softmax, cross-entropy, sampling from distributions
  • Command line β€” terminal, pip/conda, environment variables

πŸ“ Changelog

Date Changes
April 2026 Added Appendix C β€” HuggingFace Model Config Field Reference (LLM, VLM, VLA configs, kernel pitfall checklist)
April 2026 Added Ch 9 (Data), Ch 11 (Optimizers & Loss Functions), Ch 14 (PEFT Deep Dive), Ch 23 (Knowledge Distillation & QAD) β€” renumbered all chapters accordingly
April 2026 Initial release β€” 34 chapters + 2 appendices

Last updated: April 2026