Appendix A β A Brief History of Language Modelling
βThose who cannot remember the past are condemned to reimplement it.β
Timeline
The Arc of Language Modelling
1948
Claude Shannon β "A Mathematical Theory of Communication". Entropy of English, N-gram prediction, Markov chains.
1980s
N-gram language models with smoothing (Katz, Kneser-Ney). Dominated NLP for two decades.
2003
Bengio et al. β "A Neural Probabilistic Language Model". First neural LM: learn word embeddings + predict next word with a feedforward network.
2011
RNN language models (Mikolov). Recurrent networks for variable-length sequences.
2013
Word2Vec (Mikolov et al.). Efficient word embeddings via skip-gram and CBOW. "King β Man + Woman = Queen".
2014
Seq2Seq (Sutskever et al.). Encoder-decoder RNNs for machine translation.
2015
Attention mechanism (Bahdanau et al.). Allow decoder to "attend" to relevant encoder states.
2017
Transformer (Vaswani et al.) β "Attention Is All You Need". Self-attention replaces recurrence entirely. Parallelizable, scalable.
2018
GPT-1 (Radford et al.) β decoder-only transformer, 117M params, pre-train then fine-tune. BERT (Devlin et al.) β bidirectional encoder, MLM pre-training. Revolutionized NLU.
2019
GPT-2 (1.5B) β "Language Models are Unsupervised Multitask Learners". Zero-shot task performance. T5 β text-to-text framework.
2020
GPT-3 (175B) β in-context learning emerges at scale. Few-shot prompting. Kaplan scaling laws.
2022
ChatGPT β RLHF alignment makes LLMs usable as assistants. Chinchilla scaling laws. InstructGPT. PaLM (540B).
2023
GPT-4, LLaMA (open weights revolution), Mistral-7B, Mamba (SSMs), Flash Attention 2.
2024
LLaMA-3, DeepSeek-V2/V3, Qwen-2.5, o1 (reasoning), Claude 3.5, Gemini 1.5 (2M context). MoE at scale.
2025
DeepSeek-R1 (open reasoning), reasoning models proliferate, agents become practical, efficiency frontier pushed (small models, big data).
Key Papers
| Year | Paper | Authors | Contribution |
|---|---|---|---|
| 1948 | A Mathematical Theory of Communication | Shannon | Information theory, entropy of language |
| 2003 | A Neural Probabilistic Language Model | Bengio et al. | First neural language model |
| 2013 | Efficient Estimation of Word Representations | Mikolov et al. | Word2Vec β skip-gram & CBOW |
| 2014 | Sequence to Sequence Learning | Sutskever et al. | Encoder-decoder for translation |
| 2015 | Neural Machine Translation by Jointly Learning to Align and Translate | Bahdanau et al. | Attention mechanism |
| 2017 | Attention Is All You Need | Vaswani et al. | The Transformer |
| 2018 | Improving Language Understanding by Generative Pre-Training | Radford et al. | GPT-1 |
| 2018 | BERT: Pre-training of Deep Bidirectional Transformers | Devlin et al. | BERT, MLM |
| 2019 | Language Models are Unsupervised Multitask Learners | Radford et al. | GPT-2 |
| 2020 | Language Models are Few-Shot Learners | Brown et al. | GPT-3, in-context learning |
| 2020 | Scaling Laws for Neural Language Models | Kaplan et al. | Power-law scaling |
| 2021 | LoRA: Low-Rank Adaptation of Large Language Models | Hu et al. | Parameter-efficient fine-tuning |
| 2022 | Training Compute-Optimal Large Language Models | Hoffmann et al. | Chinchilla scaling laws |
| 2022 | Training language models to follow instructions (InstructGPT) | Ouyang et al. | RLHF for alignment |
| 2022 | Chain-of-Thought Prompting | Wei et al. | Step-by-step reasoning |
| 2023 | LLaMA: Open and Efficient Foundation Language Models | Touvron et al. | Open weights |
| 2023 | FlashAttention-2 | Dao | IO-aware exact attention |
| 2023 | Direct Preference Optimization | Rafailov et al. | DPO β simpler alternative to RLHF |
| 2023 | Mamba: Linear-Time Sequence Modeling | Gu & Dao | Selective state spaces |
| 2024 | The Llama 3 Herd of Models | Meta AI | LLaMA-3 family, 8Bβ405B |
| 2024 | DeepSeek-V3 Technical Report | DeepSeek | 671B MoE, FP8 training |
| 2025 | DeepSeek-R1 | DeepSeek | RL-trained reasoning, open weights |
The N-gram Era
Before neural networks, language models counted word sequences:
\[P(w_n | w_1, \ldots, w_{n-1}) \approx P(w_n | w_{n-k}, \ldots, w_{n-1})\]1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
# Simple bigram model in NumPy
import numpy as np
from collections import Counter
def train_bigram(corpus):
"""Train a bigram language model from a list of sentences."""
bigram_counts = Counter()
unigram_counts = Counter()
for sentence in corpus:
tokens = ['<s>'] + sentence.split() + ['</s>']
for i in range(len(tokens) - 1):
bigram_counts[(tokens[i], tokens[i+1])] += 1
unigram_counts[tokens[i]] += 1
# Convert to probabilities with add-1 smoothing
vocab = set(unigram_counts.keys())
V = len(vocab)
def probability(word, context):
return (bigram_counts[(context, word)] + 1) / (unigram_counts[context] + V)
return probability, vocab
def perplexity(prob_fn, vocab, test_sentence):
tokens = ['<s>'] + test_sentence.split() + ['</s>']
log_prob = sum(np.log2(prob_fn(tokens[i+1], tokens[i])) for i in range(len(tokens)-1))
return 2 ** (-log_prob / (len(tokens) - 1))
N-grams were the state of the art for speech recognition and machine translation until ~2014. Their fundamental limitation: they canβt generalize beyond seen N-grams and require exponential storage for long contexts.
The Neural Revolution
Bengioβs 2003 paper introduced two ideas that still underpin modern LLMs:
- Learned embeddings: represent words as dense vectors (not one-hot)
- Shared parameters: a neural network that generalizes across contexts
This evolved through RNNs β LSTMs β Attention β Transformers, each solving a limitation of the previous:
| Architecture | Key Limitation Solved | New Limitation |
|---|---|---|
| Feedforward (2003) | Canβt handle variable length | Fixed context window |
| RNN (2011) | Variable-length sequences | Vanishing gradients, sequential |
| LSTM (2014) | Long-range dependencies | Still sequential, slow to train |
| Seq2Seq + Attention (2015) | Encoder bottleneck | Still recurrent |
| Transformer (2017) | Fully parallel training | Quadratic attention cost |
See the main chapters for deep dives: Chapter 3 β The Transformer, Chapter 4 β Attention.
β Back to Table of Contents Β· Appendix B β Tokenization Deep Dive β
Last updated: April 2026