← Back to Table of Contents

Appendix A β€” A Brief History of Language Modelling

β€œThose who cannot remember the past are condemned to reimplement it.”

Timeline

The Arc of Language Modelling
1948
Claude Shannon β€” "A Mathematical Theory of Communication". Entropy of English, N-gram prediction, Markov chains.
1980s
N-gram language models with smoothing (Katz, Kneser-Ney). Dominated NLP for two decades.
2003
Bengio et al. β€” "A Neural Probabilistic Language Model". First neural LM: learn word embeddings + predict next word with a feedforward network.
2011
RNN language models (Mikolov). Recurrent networks for variable-length sequences.
2013
Word2Vec (Mikolov et al.). Efficient word embeddings via skip-gram and CBOW. "King βˆ’ Man + Woman = Queen".
2014
Seq2Seq (Sutskever et al.). Encoder-decoder RNNs for machine translation.
2015
Attention mechanism (Bahdanau et al.). Allow decoder to "attend" to relevant encoder states.
2017
Transformer (Vaswani et al.) β€” "Attention Is All You Need". Self-attention replaces recurrence entirely. Parallelizable, scalable.
2018
GPT-1 (Radford et al.) β€” decoder-only transformer, 117M params, pre-train then fine-tune. BERT (Devlin et al.) β€” bidirectional encoder, MLM pre-training. Revolutionized NLU.
2019
GPT-2 (1.5B) β€” "Language Models are Unsupervised Multitask Learners". Zero-shot task performance. T5 β€” text-to-text framework.
2020
GPT-3 (175B) β€” in-context learning emerges at scale. Few-shot prompting. Kaplan scaling laws.
2022
ChatGPT β€” RLHF alignment makes LLMs usable as assistants. Chinchilla scaling laws. InstructGPT. PaLM (540B).
2023
GPT-4, LLaMA (open weights revolution), Mistral-7B, Mamba (SSMs), Flash Attention 2.
2024
LLaMA-3, DeepSeek-V2/V3, Qwen-2.5, o1 (reasoning), Claude 3.5, Gemini 1.5 (2M context). MoE at scale.
2025
DeepSeek-R1 (open reasoning), reasoning models proliferate, agents become practical, efficiency frontier pushed (small models, big data).

Key Papers

Year Paper Authors Contribution
1948 A Mathematical Theory of Communication Shannon Information theory, entropy of language
2003 A Neural Probabilistic Language Model Bengio et al. First neural language model
2013 Efficient Estimation of Word Representations Mikolov et al. Word2Vec β€” skip-gram & CBOW
2014 Sequence to Sequence Learning Sutskever et al. Encoder-decoder for translation
2015 Neural Machine Translation by Jointly Learning to Align and Translate Bahdanau et al. Attention mechanism
2017 Attention Is All You Need Vaswani et al. The Transformer
2018 Improving Language Understanding by Generative Pre-Training Radford et al. GPT-1
2018 BERT: Pre-training of Deep Bidirectional Transformers Devlin et al. BERT, MLM
2019 Language Models are Unsupervised Multitask Learners Radford et al. GPT-2
2020 Language Models are Few-Shot Learners Brown et al. GPT-3, in-context learning
2020 Scaling Laws for Neural Language Models Kaplan et al. Power-law scaling
2021 LoRA: Low-Rank Adaptation of Large Language Models Hu et al. Parameter-efficient fine-tuning
2022 Training Compute-Optimal Large Language Models Hoffmann et al. Chinchilla scaling laws
2022 Training language models to follow instructions (InstructGPT) Ouyang et al. RLHF for alignment
2022 Chain-of-Thought Prompting Wei et al. Step-by-step reasoning
2023 LLaMA: Open and Efficient Foundation Language Models Touvron et al. Open weights
2023 FlashAttention-2 Dao IO-aware exact attention
2023 Direct Preference Optimization Rafailov et al. DPO β€” simpler alternative to RLHF
2023 Mamba: Linear-Time Sequence Modeling Gu & Dao Selective state spaces
2024 The Llama 3 Herd of Models Meta AI LLaMA-3 family, 8B–405B
2024 DeepSeek-V3 Technical Report DeepSeek 671B MoE, FP8 training
2025 DeepSeek-R1 DeepSeek RL-trained reasoning, open weights

The N-gram Era

Before neural networks, language models counted word sequences:

\[P(w_n | w_1, \ldots, w_{n-1}) \approx P(w_n | w_{n-k}, \ldots, w_{n-1})\]
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
# Simple bigram model in NumPy
import numpy as np
from collections import Counter

def train_bigram(corpus):
    """Train a bigram language model from a list of sentences."""
    bigram_counts = Counter()
    unigram_counts = Counter()
    
    for sentence in corpus:
        tokens = ['<s>'] + sentence.split() + ['</s>']
        for i in range(len(tokens) - 1):
            bigram_counts[(tokens[i], tokens[i+1])] += 1
            unigram_counts[tokens[i]] += 1
    
    # Convert to probabilities with add-1 smoothing
    vocab = set(unigram_counts.keys())
    V = len(vocab)
    
    def probability(word, context):
        return (bigram_counts[(context, word)] + 1) / (unigram_counts[context] + V)
    
    return probability, vocab

def perplexity(prob_fn, vocab, test_sentence):
    tokens = ['<s>'] + test_sentence.split() + ['</s>']
    log_prob = sum(np.log2(prob_fn(tokens[i+1], tokens[i])) for i in range(len(tokens)-1))
    return 2 ** (-log_prob / (len(tokens) - 1))

N-grams were the state of the art for speech recognition and machine translation until ~2014. Their fundamental limitation: they can’t generalize beyond seen N-grams and require exponential storage for long contexts.

The Neural Revolution

Bengio’s 2003 paper introduced two ideas that still underpin modern LLMs:

  1. Learned embeddings: represent words as dense vectors (not one-hot)
  2. Shared parameters: a neural network that generalizes across contexts

This evolved through RNNs β†’ LSTMs β†’ Attention β†’ Transformers, each solving a limitation of the previous:

Architecture Key Limitation Solved New Limitation
Feedforward (2003) Can’t handle variable length Fixed context window
RNN (2011) Variable-length sequences Vanishing gradients, sequential
LSTM (2014) Long-range dependencies Still sequential, slow to train
Seq2Seq + Attention (2015) Encoder bottleneck Still recurrent
Transformer (2017) Fully parallel training Quadratic attention cost

See the main chapters for deep dives: Chapter 3 β€” The Transformer, Chapter 4 β€” Attention.


← Back to Table of Contents Β· Appendix B β€” Tokenization Deep Dive β†’


Last updated: April 2026