← Back to Table of Contents

Chapter 1 β€” Introduction to Language Modelling

β€œThe problem of predicting the next word is, in a sense, the problem of understanding language.” β€” Yoshua Bengio

The Core Idea

A language model is a system that assigns probabilities to sequences of words. At its heart, the task is deceptively simple: given some text, predict what comes next. This single objective β€” next-token prediction β€” turns out to be powerful enough to produce systems that can write code, answer questions, translate languages, and reason about the world.

Every modern large language model (LLM), from GPT-4 to LLaMA to Gemini, is trained on some variation of this principle. The model reads a sequence of tokens and learns to predict the probability distribution over the next token:

\[P(x_{t+1} \mid x_1, x_2, \ldots, x_t)\]
The Language Modelling Objective
The cat sat
on (83%)

The model doesn’t β€œunderstand” in a human sense β€” it builds statistical representations of language so rich that understanding-like behavior emerges from the training process.

From Text to Tensors β€” The Full Pipeline

Before a model can process text, the text must become numbers. Before those numbers mean anything, they must be transformed into rich vector representations. The pipeline from raw text to a model’s prediction involves several stages, each covered in depth in later chapters:

The Language Model Pipeline
πŸ“ Raw Text "The cat sat on"
πŸ”€ Tokenizer text β†’ token IDs
πŸ“Š Embedding Layer token IDs β†’ dense vectors
🧠 Transformer Blocks self-attention + feed-forward Γ— N
πŸ“ˆ Output Head hidden states β†’ vocabulary logits
🎯 Sampling logits β†’ next token

And here are the tensor shapes at each stage β€” a pattern you’ll see throughout this guide:

Tensor Shapes Through the Pipeline
Token IDs
[ B , T ]
Embeddings
[ B , T , d_model ]
After Transformer
[ B , T , d_model ]
Logits
[ B , T , V ]

Where B is batch size, T is sequence length, d_model is the hidden dimension (e.g., 4096), and V is the vocabulary size (e.g., 32000 for LLaMA).

A Brief History

Language modelling has evolved through several paradigm shifts over six decades. For the full historical deep dive, see Appendix A. Here are the key milestones:

Key Milestones in Language Modelling
1948–1990s
Statistical Era
Shannon's information theory β†’ N-gram models β†’ smoothing techniques
2003
Neural Language Models
Bengio et al. β€” first neural network that learns word representations and predicts next words
2013
Word2Vec
Mikolov et al. β€” efficient word embeddings, "king - man + woman = queen"
2017
Transformers
Vaswani et al. β€” "Attention Is All You Need" β€” the architecture that changed everything
2018–2019
Pre-training Revolution
GPT, BERT, GPT-2 β€” large-scale pre-training on internet text
2020–2023
The Scaling Era
GPT-3, Chinchilla, LLaMA, GPT-4 β€” scaling laws, emergence, open-source explosion
2024–Present
Reasoning & Efficiency
o1, DeepSeek-R1, Mamba, MoE β€” test-time compute, efficient architectures, open weights

Tokenization Essentials

Before a model sees any text, a tokenizer breaks it into discrete units called tokens. Modern LLMs use subword tokenization β€” a middle ground between character-level and word-level that handles rare words and multiple languages efficiently.

For the complete tokenization deep dive (BPE algorithm walkthrough, training tokenizers from scratch), see Appendix B. Here’s what you need to know:

Method Used By Key Idea
BPE (Byte Pair Encoding) GPT, LLaMA, Mistral Iteratively merge most frequent byte pairs
WordPiece BERT, DistilBERT Similar to BPE but uses likelihood-based merges
Unigram T5, ALBERT Start with large vocab, prune to target size
SentencePiece LLaMA, T5 Language-agnostic, treats input as raw bytes

The tokenizer defines the vocabulary size V, which determines the final projection layer. A typical modern LLM uses V = 32,000–128,000 tokens.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3-8B")

text = "Language models predict the next token."
tokens = tokenizer.encode(text)
print(f"Text: {text}")
print(f"Token IDs: {tokens}")
print(f"Tokens: {tokenizer.convert_ids_to_tokens(tokens)}")
print(f"Vocab size: {tokenizer.vocab_size}")
# Text: Language models predict the next token.
# Token IDs: [14658, 4211, 7963, 279, 1828, 4037, 13]
# Tokens: ['Language', ' models', ' predict', ' the', ' next', ' token', '.']
# Vocab size: 128256

The Simplest Language Model

To build intuition, here is the simplest possible language model β€” a bigram model that predicts the next token based only on the current one:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
import numpy as np

# A tiny corpus
corpus = "the cat sat on the mat the cat ate the rat"
words = corpus.split()
vocab = sorted(set(words))
w2i = {w: i for i, w in enumerate(vocab)}

# Count bigram frequencies
V = len(vocab)
counts = np.zeros((V, V))
for w1, w2 in zip(words[:-1], words[1:]):
    counts[w2i[w1], w2i[w2]] += 1

# Normalize to probabilities (add-1 smoothing)
probs = (counts + 1) / (counts + 1).sum(axis=1, keepdims=True)

# Predict next word
current = "the"
next_probs = probs[w2i[current]]
for w, p in sorted(zip(vocab, next_probs), key=lambda x: -x[1]):
    print(f"  P({w} | {current}) = {p:.3f}")
# P(cat | the) = 0.273
# P(mat | the) = 0.182
# P(rat | the) = 0.182
# ...

A bigram model is limited β€” it has no memory beyond one token. A transformer-based language model considers the entire context window (thousands or millions of tokens), learns rich representations, and can capture long-range dependencies. But the core idea is the same: assign probabilities to what comes next.

How Modern LLMs Differ

Modern LLMs like GPT-4, LLaMA 3, and Gemini all share the same fundamental architecture (the transformer, Chapter 3) but differ in:

What Differentiates Modern LLMs
πŸ“
Architecture
Attention type, norm placement, FFN variant, context length
πŸ“Š
Training Data
Corpus size, quality filters, domain mix, deduplication
βš–οΈ
Scale
Parameter count, compute budget, training tokens
🎯
Alignment
RLHF, DPO, safety training, instruction tuning
πŸ”€
Tokenizer
Vocabulary size, BPE variant, multilingual support
🌐
Modality
Text-only vs multimodal (vision, audio, code)

Putting It All Together β€” 10 Lines

Here’s the entire pipeline β€” from text to generated response β€” in 10 lines of Python:

1
2
3
4
5
6
7
8
9
10
11
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model_name = "meta-llama/Llama-3-8B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16, device_map="auto")

prompt = "Explain language modelling in one sentence:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)  # [1, T]
outputs = model.generate(**inputs, max_new_tokens=50)               # [1, T + 50]
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Every component in those 10 lines β€” the tokenizer, the embedding layer, the transformer blocks, the sampling strategy β€” is a chapter in this guide.

What’s Next

With the big picture in place, the next chapter dives into embeddings β€” how token IDs become the rich vector representations that transformers actually process.

Next: Chapter 2 β€” Embeddings & Representations β†’


Last updated: April 2026