Phase 6 · Tools, Agents & the APIModule 23~36 min read

Tools & Function Calling

Let the model do things: search, run code, hit an API. Define tools, describe them well, and handle the model's calls — the foundation of every agent.

What you'll learn

A model on its own can only produce text. Tools (also called function calling) let it act: search the web, run a calculation, query a database, hit an API. You describe the tools; the model decides when to call them; your code runs them and hands back the result. This is the foundation of every agent.

By the end of this module you'll be able to:

  • Explain why tools dramatically expand what a model can do
  • Define a tool with a clear schema the model can call
  • Write tool descriptions the model actually understands
  • Run the call → execute → return loop safely

Why models need tools

Models are unreliable at exactly the things ordinary software is great at: precise arithmetic, looking up live data, remembering specifics, taking real actions. Tools bridge that gap. Instead of guessing today's exchange rate, the model calls a currency API and reads the real number — the same insight behind ReAct (Module 13), now made concrete.

Defining a tool

A tool definition is essentially a typed contract: a name, a description of what it does and when to use it, and a schema for its arguments (the same schema thinking as Module 14). The model never runs code — it just emits a structured request to call the tool with specific arguments.

tool_def.py
# You describe each tool so the model knows when and how to call it.
get_weather = {
    "name": "get_weather",
    "description": "Get the current weather for a city. "
                   "Use when the user asks about weather or temperature.",
    "parameters": {
        "type": "object",
        "properties": {
            "city": {"type": "string", "description": "City name, e.g. 'Paris'"},
            "unit": {"type": "string", "enum": ["c", "f"], "default": "c"}
        },
        "required": ["city"]
    }
}
A tool is a name, a description, and a parameter schema. Enums and 'required' fields keep the model's calls valid.

Descriptions that work

The description is a prompt — the model relies on it to decide whether and how to call the tool. Vague descriptions cause missed or wrong calls. Make them count:

  • Say what it does AND when to use it ("Use when the user asks about weather").
  • Describe every parameter with types, units, and an example value.
  • Constrain with enums so the model can't invent an invalid argument.
  • Note limits — what the tool can't do, so the model doesn't misuse it.

Key idea

Treat tool descriptions as first-class prompt engineering. Most tool-use failures aren't model failures — they're under-specified descriptions. If the model calls the wrong tool or bad arguments, fix the wording before anything else.

Call, execute, return

Tool use is a loop, and a crucial point: the model never executes anything — you do. It only requests a call; your code runs the real function and returns the result. The model then continues with that result in hand, possibly calling more tools before it answers.

The tool-use loop

1 · Model decides

It sees the tools and requests one, with arguments.

2 · Your code runs it

You execute the real function and get a result.

3 · Result goes back

You return the output to the model as a new message.

4 · Model continues

It uses the result to answer — or calls another tool.

↻ the model can loop through tools before answering

The model requests a tool; your code executes it and returns the result; the model continues. It can loop through several tools.
tool_loop.py
# The tool-use loop.
messages = [{"role": "user", "content": "What's the weather in Paris?"}]

while True:
    resp = client.messages.create(model=M, messages=messages, tools=[get_weather])

    if resp.tool_call:                       # 1. model asked for a tool
        args = resp.tool_call.arguments      #    {"city": "Paris"}
        result = get_weather_impl(**args)    # 2. YOUR code runs the real function
        messages.append(resp.message)        # 3. record the request...
        messages.append({"role": "tool",     #    ...and return the result
                         "content": result})
        continue                             # 4. loop: let the model use it
    else:
        print(resp.output_text)              # done - final answer
        break
Your code owns execution. The model asks; you run the function and feed the result back; repeat until it answers.

Note

Models can also request several tool calls at once (parallel) or chain them across turns. The loop is the same — run each requested call, return each result, and let the model proceed.

Errors & tool safety

Because tools take real actions, they're where prompting meets real-world risk. Handle it deliberately:

  • Return errors as data. If a tool fails, hand the model a clear error message so it can retry or adjust, rather than crashing.
  • Validate arguments before executing — the model can produce out-of-range or malformed inputs.
  • Least privilege. Give the model only the tools it needs, scoped as narrowly as possible.
  • Guard destructive actions. Require confirmation for anything that deletes, sends, pays, or can't be undone.

Watch out

Tools plus untrusted input is the highest-risk combination in all of prompting: a malicious document could try to make the model call a tool it shouldn't. We tackle this head-on in Module 28 (Prompt Injection) — never wire a powerful tool to untrusted input without safeguards.

Recap & quick check

Key takeaways

  • Tools let a model act — search, compute, look up, take actions — covering exactly what models are unreliable at.
  • A tool definition is a typed contract: name, a description of what/when, and a parameter schema (use enums + required).
  • Descriptions are prompts; most tool-use failures are under-specified descriptions, not model failures.
  • The model only requests calls — your code executes them and returns results; the loop repeats until it answers.
  • Handle tool errors as data, validate arguments, grant least privilege, and gate destructive or irreversible actions.

Quick check

1. In tool use / function calling, who actually executes the function?

2. Why do tools make models much more capable?

3. A model keeps calling the wrong tool. What should you fix first?

4. Which is a key safety practice for tools?

Give a model tools and a goal, and let it loop on its own, and you have an agent. That's next. Next up: Module 24 — Building Agents: The Reasoning Loop.