← Back to Table of Contents

Chapter 35 β€” Reasoning Models

β€œChain-of-thought doesn’t teach the model new facts β€” it teaches the model to use the facts it already knows, step by step.”

The Reasoning Gap

Standard LLMs generate token by token, choosing the most likely next token. This works well for fluent text but struggles with tasks that require multi-step reasoning:

  • β€œWhat is 347 Γ— 892?” β€” requires carrying digits across steps
  • β€œIf all roses are flowers and some flowers fade quickly, can we conclude some roses fade quickly?” β€” requires logical chaining
  • β€œWrite a function that finds the longest increasing subsequence” β€” requires algorithm design

Reasoning models address this by generating intermediate reasoning steps before the final answer.

Chain-of-Thought (CoT) Prompting

Chain-of-thought prompting comparison versus standard prompting, plus test-time compute scaling with Best-of-N, PRM, and MCTS strategies
CoT makes the model reason step-by-step; test-time compute scaling (o1-style) extends this with search over reasoning paths

Wei et al. (2022) showed that adding β€œLet’s think step by step” dramatically improves math and logic:

Standard vs Chain-of-Thought
Standard Prompting
  • Q: Roger has 5 balls. He buys 2 cans of 3 balls each. How many does he have?
  • A: 11 βœ—
  • (No reasoning shown, just guesses)
CoT Prompting
  • Q: ...same question...
  • A: Roger started with 5 balls. He bought 2 cans Γ— 3 balls = 6 balls. Total: 5 + 6 = 11. βœ“
  • (Step-by-step reasoning leads to correct answer)
1
2
3
4
5
6
7
8
9
10
11
12
13
# Zero-shot CoT β€” just append the magic phrase
prompt = """Q: A store had 45 apples. They sold 12 on Monday and received 
a shipment of 30 on Tuesday. Then sold 25 on Wednesday. How many remain?

Let's think step by step:"""

# Few-shot CoT β€” provide worked examples in the prompt
few_shot_prompt = """Q: If there are 3 cars in the parking lot and 2 more arrive, 
how many cars are there?
A: There are originally 3 cars. 2 more arrive. 3 + 2 = 5. The answer is 5.

Q: {actual_question}
A:"""

CoT works because it:

  1. Decomposes complex problems into simpler sub-problems
  2. Allocates more compute: more tokens = more FLOPs per problem
  3. Creates intermediate representations the model can condition on
  4. Only effective at β‰₯ ~60B parameters (smaller models generate incoherent chains)

Self-Consistency

Sample multiple chain-of-thought reasoning paths and take the majority vote:

Self-Consistency (Wang et al., 2022)
Same question asked N times with temperature > 0
↓
Path 1: ... 5 + 6 = 11 β†’ Answer: 11
Path 2: ... 5 + 6 = 11 β†’ Answer: 11
Path 3: ... 5 + 3 = 8 β†’ Answer: 8 (wrong reasoning)
Path 4: ... 5 + 6 = 11 β†’ Answer: 11
Path 5: ... 5 + 6 = 12 β†’ Answer: 12 (arithmetic error)
↓
Majority vote: 11 wins (3/5) β†’ Final answer: 11 βœ“
1
2
3
4
5
6
7
8
9
10
11
12
13
import collections

def self_consistency(model, prompt, n_samples=10, temperature=0.7):
    """Generate multiple reasoning paths and majority vote."""
    answers = []
    for _ in range(n_samples):
        response = model.generate(prompt, temperature=temperature, max_tokens=512)
        answer = extract_final_answer(response)  # parse the numeric/categorical answer
        answers.append(answer)
    
    # Majority vote
    counter = collections.Counter(answers)
    return counter.most_common(1)[0][0]

Self-consistency improves GSM8K accuracy by 5–15% over single-sample CoT.

Test-Time Compute Scaling (o1-style)

The breakthrough insight: scaling compute at inference time (more thinking tokens) can be as effective as scaling model size:

Two Dimensions of Scaling
Pre-Training Scaling
  • More parameters, more data
  • Fixed compute per token at inference
  • Expensive to scale (months of GPU time)
  • Improves general knowledge + capabilities
Test-Time Scaling
  • Same model, more inference tokens
  • Variable compute per problem (think harder on hard problems)
  • Cheap to scale (just generate more tokens)
  • Improves reasoning on specific problems

How o1 Works (Conceptual)

OpenAI’s o1 family generates long internal reasoning chains before answering:

o1 Reasoning Process
User question received
↓
[Internal thinking β€” not shown to user]
Break problem into sub-problems
Try approach A β†’ hit dead end β†’ backtrack
Try approach B β†’ partial progress
Verify intermediate results
Complete the solution
Self-check: does the answer make sense?
↓
[Summary answer shown to user]

Key capabilities trained into reasoning models:

  • Backtracking: β€œWait, that’s wrong. Let me try a different approach.”
  • Verification: β€œLet me check: 347 Γ— 892 = 309,524. Verify: 892 Γ— 300 = 267,600, 892 Γ— 47 = 41,924. Sum: 309,524. βœ“β€
  • Self-reflection: β€œThis approach is getting too complicated. There might be a simpler way.”

Process Reward Models vs Outcome Reward Models

Reward Model Types
Outcome Reward Model (ORM)
  • Scores only the final answer
  • Correct answer β†’ reward = 1, wrong β†’ 0
  • Sparse signal β€” hard to learn from
  • Can't distinguish lucky guesses from good reasoning
Process Reward Model (PRM)
  • Scores each reasoning step independently
  • Step 1: correct βœ“, Step 2: correct βœ“, Step 3: error βœ—
  • Dense signal β€” model learns where it went wrong
  • Used in o1 and similar systems

Training a PRM requires step-level annotations (expensive) or automated verification (for math/code where answers can be checked programmatically).

Open-Source Reasoning Models

Model Approach Key Innovation
DeepSeek-R1 RL-trained reasoning Long CoT via RL, open weights, distillation to smaller models
QwQ (Qwen) Long reasoning chains Extended thinking with self-reflection
Marco-o1 Open replication of o1-style Process reward + MCTS search
s1 (simple test-time scaling) Budget forcing Control reasoning length via β€œWait” token injection

DeepSeek-R1

DeepSeek-R1 demonstrated that reasoning can emerge from pure RL without supervised CoT data:

  1. Start with DeepSeek-V3 base model
  2. Apply RL (GRPO) with only outcome verification (correct/incorrect)
  3. The model spontaneously learns to reason, self-verify, and backtrack
  4. Distill the reasoning ability into smaller models (1.5B, 7B, 14B, 32B, 70B)

Budget Forcing: Controlling Reasoning Depth

Not every question needs deep reasoning. Budget forcing lets you control how much the model β€œthinks”:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
# Conceptual: control reasoning depth
def generate_with_budget(model, prompt, thinking_budget="medium"):
    """
    thinking_budget: "low" (fast), "medium", "high" (thorough)
    """
    budgets = {"low": 256, "medium": 1024, "high": 4096}
    max_thinking_tokens = budgets[thinking_budget]
    
    response = model.generate(
        prompt,
        max_tokens=max_thinking_tokens + 512,  # thinking + answer
        stop=["</answer>"],
    )
    return extract_answer(response)

The Reasoning Landscape

Evolution of LLM Reasoning
2022
Chain-of-Thought prompting (Wei et al.) β€” "Let's think step by step"
2022
Self-Consistency (Wang et al.) β€” majority vote over multiple CoT paths
2023
Tree of Thoughts β€” search over reasoning trees with LLM-guided evaluation
2024
o1 (OpenAI) β€” RL-trained internal reasoning with backtracking + verification
2025
DeepSeek-R1, QwQ β€” open-source reasoning models, RL-based emergence of CoT

What’s Next

Reasoning helps models think better with their internal knowledge. But what if the model’s knowledge is outdated or insufficient? The next chapter covers Retrieval-Augmented Generation (RAG) β€” grounding model outputs in external, up-to-date information.

← Previous: Chapter 34 β€” SSMs & Beyond Transformers Β· Next: Chapter 36 β€” Retrieval-Augmented Generation β†’


Last updated: April 2026