Chapter 35 β Reasoning Models
βChain-of-thought doesnβt teach the model new facts β it teaches the model to use the facts it already knows, step by step.β
The Reasoning Gap
Standard LLMs generate token by token, choosing the most likely next token. This works well for fluent text but struggles with tasks that require multi-step reasoning:
- βWhat is 347 Γ 892?β β requires carrying digits across steps
- βIf all roses are flowers and some flowers fade quickly, can we conclude some roses fade quickly?β β requires logical chaining
- βWrite a function that finds the longest increasing subsequenceβ β requires algorithm design
Reasoning models address this by generating intermediate reasoning steps before the final answer.
Chain-of-Thought (CoT) Prompting
Wei et al. (2022) showed that adding βLetβs think step by stepβ dramatically improves math and logic:
- Q: Roger has 5 balls. He buys 2 cans of 3 balls each. How many does he have?
- A: 11 β
- (No reasoning shown, just guesses)
- Q: ...same question...
- A: Roger started with 5 balls. He bought 2 cans Γ 3 balls = 6 balls. Total: 5 + 6 = 11. β
- (Step-by-step reasoning leads to correct answer)
1
2
3
4
5
6
7
8
9
10
11
12
13
# Zero-shot CoT β just append the magic phrase
prompt = """Q: A store had 45 apples. They sold 12 on Monday and received
a shipment of 30 on Tuesday. Then sold 25 on Wednesday. How many remain?
Let's think step by step:"""
# Few-shot CoT β provide worked examples in the prompt
few_shot_prompt = """Q: If there are 3 cars in the parking lot and 2 more arrive,
how many cars are there?
A: There are originally 3 cars. 2 more arrive. 3 + 2 = 5. The answer is 5.
Q: {actual_question}
A:"""
CoT works because it:
- Decomposes complex problems into simpler sub-problems
- Allocates more compute: more tokens = more FLOPs per problem
- Creates intermediate representations the model can condition on
- Only effective at β₯ ~60B parameters (smaller models generate incoherent chains)
Self-Consistency
Sample multiple chain-of-thought reasoning paths and take the majority vote:
1
2
3
4
5
6
7
8
9
10
11
12
13
import collections
def self_consistency(model, prompt, n_samples=10, temperature=0.7):
"""Generate multiple reasoning paths and majority vote."""
answers = []
for _ in range(n_samples):
response = model.generate(prompt, temperature=temperature, max_tokens=512)
answer = extract_final_answer(response) # parse the numeric/categorical answer
answers.append(answer)
# Majority vote
counter = collections.Counter(answers)
return counter.most_common(1)[0][0]
Self-consistency improves GSM8K accuracy by 5β15% over single-sample CoT.
Test-Time Compute Scaling (o1-style)
The breakthrough insight: scaling compute at inference time (more thinking tokens) can be as effective as scaling model size:
- More parameters, more data
- Fixed compute per token at inference
- Expensive to scale (months of GPU time)
- Improves general knowledge + capabilities
- Same model, more inference tokens
- Variable compute per problem (think harder on hard problems)
- Cheap to scale (just generate more tokens)
- Improves reasoning on specific problems
How o1 Works (Conceptual)
OpenAIβs o1 family generates long internal reasoning chains before answering:
Key capabilities trained into reasoning models:
- Backtracking: βWait, thatβs wrong. Let me try a different approach.β
- Verification: βLet me check: 347 Γ 892 = 309,524. Verify: 892 Γ 300 = 267,600, 892 Γ 47 = 41,924. Sum: 309,524. ββ
- Self-reflection: βThis approach is getting too complicated. There might be a simpler way.β
Process Reward Models vs Outcome Reward Models
- Scores only the final answer
- Correct answer β reward = 1, wrong β 0
- Sparse signal β hard to learn from
- Can't distinguish lucky guesses from good reasoning
- Scores each reasoning step independently
- Step 1: correct β, Step 2: correct β, Step 3: error β
- Dense signal β model learns where it went wrong
- Used in o1 and similar systems
Training a PRM requires step-level annotations (expensive) or automated verification (for math/code where answers can be checked programmatically).
Open-Source Reasoning Models
| Model | Approach | Key Innovation |
|---|---|---|
| DeepSeek-R1 | RL-trained reasoning | Long CoT via RL, open weights, distillation to smaller models |
| QwQ (Qwen) | Long reasoning chains | Extended thinking with self-reflection |
| Marco-o1 | Open replication of o1-style | Process reward + MCTS search |
| s1 (simple test-time scaling) | Budget forcing | Control reasoning length via βWaitβ token injection |
DeepSeek-R1
DeepSeek-R1 demonstrated that reasoning can emerge from pure RL without supervised CoT data:
- Start with DeepSeek-V3 base model
- Apply RL (GRPO) with only outcome verification (correct/incorrect)
- The model spontaneously learns to reason, self-verify, and backtrack
- Distill the reasoning ability into smaller models (1.5B, 7B, 14B, 32B, 70B)
Budget Forcing: Controlling Reasoning Depth
Not every question needs deep reasoning. Budget forcing lets you control how much the model βthinksβ:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
# Conceptual: control reasoning depth
def generate_with_budget(model, prompt, thinking_budget="medium"):
"""
thinking_budget: "low" (fast), "medium", "high" (thorough)
"""
budgets = {"low": 256, "medium": 1024, "high": 4096}
max_thinking_tokens = budgets[thinking_budget]
response = model.generate(
prompt,
max_tokens=max_thinking_tokens + 512, # thinking + answer
stop=["</answer>"],
)
return extract_answer(response)
The Reasoning Landscape
Whatβs Next
Reasoning helps models think better with their internal knowledge. But what if the modelβs knowledge is outdated or insufficient? The next chapter covers Retrieval-Augmented Generation (RAG) β grounding model outputs in external, up-to-date information.
β Previous: Chapter 34 β SSMs & Beyond Transformers Β· Next: Chapter 36 β Retrieval-Augmented Generation β
Last updated: April 2026