Chapter 1 β Introduction to Language Modelling
βThe problem of predicting the next word is, in a sense, the problem of understanding language.β β Yoshua Bengio
The Core Idea
A language model is a system that assigns probabilities to sequences of words. At its heart, the task is deceptively simple: given some text, predict what comes next. This single objective β next-token prediction β turns out to be powerful enough to produce systems that can write code, answer questions, translate languages, and reason about the world.
Every modern large language model (LLM), from GPT-4 to LLaMA to Gemini, is trained on some variation of this principle. The model reads a sequence of tokens and learns to predict the probability distribution over the next token:
\[P(x_{t+1} \mid x_1, x_2, \ldots, x_t)\]The model doesnβt βunderstandβ in a human sense β it builds statistical representations of language so rich that understanding-like behavior emerges from the training process.
From Text to Tensors β The Full Pipeline
Before a model can process text, the text must become numbers. Before those numbers mean anything, they must be transformed into rich vector representations. The pipeline from raw text to a modelβs prediction involves several stages, each covered in depth in later chapters:
And here are the tensor shapes at each stage β a pattern youβll see throughout this guide:
Where B is batch size, T is sequence length, d_model is the hidden dimension (e.g., 4096), and V is the vocabulary size (e.g., 32000 for LLaMA).
A Brief History
Language modelling has evolved through several paradigm shifts over six decades. For the full historical deep dive, see Appendix A. Here are the key milestones:
Tokenization Essentials
Before a model sees any text, a tokenizer breaks it into discrete units called tokens. Modern LLMs use subword tokenization β a middle ground between character-level and word-level that handles rare words and multiple languages efficiently.
For the complete tokenization deep dive (BPE algorithm walkthrough, training tokenizers from scratch), see Appendix B. Hereβs what you need to know:
| Method | Used By | Key Idea |
|---|---|---|
| BPE (Byte Pair Encoding) | GPT, LLaMA, Mistral | Iteratively merge most frequent byte pairs |
| WordPiece | BERT, DistilBERT | Similar to BPE but uses likelihood-based merges |
| Unigram | T5, ALBERT | Start with large vocab, prune to target size |
| SentencePiece | LLaMA, T5 | Language-agnostic, treats input as raw bytes |
The tokenizer defines the vocabulary size V, which determines the final projection layer. A typical modern LLM uses V = 32,000β128,000 tokens.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3-8B")
text = "Language models predict the next token."
tokens = tokenizer.encode(text)
print(f"Text: {text}")
print(f"Token IDs: {tokens}")
print(f"Tokens: {tokenizer.convert_ids_to_tokens(tokens)}")
print(f"Vocab size: {tokenizer.vocab_size}")
# Text: Language models predict the next token.
# Token IDs: [14658, 4211, 7963, 279, 1828, 4037, 13]
# Tokens: ['Language', ' models', ' predict', ' the', ' next', ' token', '.']
# Vocab size: 128256
The Simplest Language Model
To build intuition, here is the simplest possible language model β a bigram model that predicts the next token based only on the current one:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
import numpy as np
# A tiny corpus
corpus = "the cat sat on the mat the cat ate the rat"
words = corpus.split()
vocab = sorted(set(words))
w2i = {w: i for i, w in enumerate(vocab)}
# Count bigram frequencies
V = len(vocab)
counts = np.zeros((V, V))
for w1, w2 in zip(words[:-1], words[1:]):
counts[w2i[w1], w2i[w2]] += 1
# Normalize to probabilities (add-1 smoothing)
probs = (counts + 1) / (counts + 1).sum(axis=1, keepdims=True)
# Predict next word
current = "the"
next_probs = probs[w2i[current]]
for w, p in sorted(zip(vocab, next_probs), key=lambda x: -x[1]):
print(f" P({w} | {current}) = {p:.3f}")
# P(cat | the) = 0.273
# P(mat | the) = 0.182
# P(rat | the) = 0.182
# ...
A bigram model is limited β it has no memory beyond one token. A transformer-based language model considers the entire context window (thousands or millions of tokens), learns rich representations, and can capture long-range dependencies. But the core idea is the same: assign probabilities to what comes next.
How Modern LLMs Differ
Modern LLMs like GPT-4, LLaMA 3, and Gemini all share the same fundamental architecture (the transformer, Chapter 3) but differ in:
Putting It All Together β 10 Lines
Hereβs the entire pipeline β from text to generated response β in 10 lines of Python:
1
2
3
4
5
6
7
8
9
10
11
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_name = "meta-llama/Llama-3-8B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16, device_map="auto")
prompt = "Explain language modelling in one sentence:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device) # [1, T]
outputs = model.generate(**inputs, max_new_tokens=50) # [1, T + 50]
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Every component in those 10 lines β the tokenizer, the embedding layer, the transformer blocks, the sampling strategy β is a chapter in this guide.
Whatβs Next
With the big picture in place, the next chapter dives into embeddings β how token IDs become the rich vector representations that transformers actually process.
Next: Chapter 2 β Embeddings & Representations β
Last updated: April 2026