← Back to Table of Contents

Chapter 37 β€” Agents & Tool Use

β€œA language model that can only generate text is a brain in a jar. An agent is a brain with hands.”

For a comprehensive deep dive into AI agents, multi-agent systems, and frameworks, see the companion guide: Agents Guide.

The Agent Loop

An agent is an LLM that can observe, reason, act, and observe again in a loop:

Agent Execution Loop
πŸ€” Reason β€” analyze the task and decide what to do next
β†’
πŸ› οΈ Act β€” call a tool (search, code execution, API call)
β†’
πŸ‘€ Observe β€” process the tool's output / result
β†’
πŸ”„ Repeat β€” until the task is complete or max steps reached

Function Calling

The foundation of tool use β€” LLMs generate structured function calls instead of (or alongside) natural text:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
# OpenAI function calling
import openai

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather for a location",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string", "description": "City name"},
                    "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
                },
                "required": ["location"],
            },
        },
    }
]

response = openai.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
    tools=tools,
    tool_choice="auto",
)

# Model responds with:
# tool_calls: [{"function": {"name": "get_weather", "arguments": '{"location": "Tokyo"}'}}]

The model doesn’t execute the function β€” it generates a structured request. Your code executes it and feeds the result back:

Function Calling Flow
User: "Weather in Tokyo?"
β†’
LLM: call get_weather("Tokyo")
β†’
System: executes function β†’ "22Β°C, sunny"
β†’
LLM: "It's 22Β°C and sunny in Tokyo."

Model Context Protocol (MCP)

MCP (Anthropic, 2024) standardizes how LLMs connect to external tools and data sources:

MCP Architecture
LLM Application (Claude, VS Code, custom app) β€” MCP Client
MCP Protocol β€” standardized JSON-RPC over stdio/SSE
MCP Server: GitHub β€” repos, issues, PRs
MCP Server: Database β€” query, schema inspection
MCP Server: File System β€” read, write, search
MCP Server: Web Search β€” Brave, Google, etc.

MCP provides a universal interface so that:

  • Any LLM application can connect to any MCP server
  • Tools are described once and discovered automatically
  • Context (resources) can be exposed alongside tools

Coding Agents

Agents that write, debug, and modify code:

Agent Approach Key Feature
GitHub Copilot IDE-integrated, multi-model Code completion, chat, workspace agents, MCP
Cursor IDE-integrated Codebase-aware, multi-file edits
Codex (OpenAI) Cloud agent Sandboxed environment, async task execution
SWE-Agent Research agent Interacts with repos via terminal commands
Devin Full IDE agent Plans, implements, tests autonomously
Aider Terminal-based Git-integrated, multi-file editing

Computer Use Agents

Agents that interact with graphical interfaces:

Computer Use Agent Loop
Take screenshot of the screen
↓
VLM analyzes the screenshot β€” identifies UI elements
↓
Decide action: click(x, y), type("text"), scroll, key("Enter")
↓
Execute action on the computer
↓
Take new screenshot β†’ repeat

Examples: Claude Computer Use, OpenAI Operator, browser-use agents.

ReAct Pattern

ReAct (Reasoning + Acting) interleaves thinking and tool use:

1
2
3
4
5
6
7
8
9
10
11
12
Question: What is the population of the capital of France?

Thought: I need to find the capital of France, then its population.
Action: search("capital of France")
Observation: The capital of France is Paris.

Thought: Now I need the population of Paris.
Action: search("population of Paris 2024")
Observation: The population of Paris is approximately 2.1 million (city proper).

Thought: I have the answer.
Answer: The population of Paris, the capital of France, is approximately 2.1 million.

Multi-Agent Systems

Complex tasks can be decomposed across specialized agents:

Multi-Agent Architecture
Orchestrator
Plans the task, delegates to specialists, synthesizes results.
Researcher
Searches the web, reads documents, gathers information.
Coder
Writes code, runs tests, fixes bugs.
Reviewer
Checks outputs for correctness, suggests improvements.

Frameworks: LangGraph, CrewAI, AutoGen, Semantic Kernel, OpenAI Swarm.

Agent Challenges

Challenge Description
Reliability Agents make mistakes that compound across steps. Error recovery is hard.
Cost Long agent traces consume many tokens. A 50-step agent run costs 50Γ— a single query.
Safety Agents can take irreversible actions (delete files, send emails). Sandboxing is critical.
Evaluation Hard to benchmark β€” success depends on environment, not just model quality.
Context window Long traces may exceed context limits. Summarization/compression needed.

What’s Next

We’ve now covered the full stack: from tokenization to training, inference, hardware, and applications. The final chapter looks at the frontier β€” where the field is heading and what remains unsolved.

← Previous: Chapter 36 β€” Retrieval-Augmented Generation Β· Next: Chapter 38 β€” The Frontier β†’


Last updated: April 2026