Agent infrastructure

Platform integration

Claude tool calling: a step-by-step guide to shipping it in production

Claude tool calling in production: a seven-step guide covering schema design, the Messages API loop, parallel calls, error handling, and common pitfalls.

7 minute read
Decorative imagery showcasing Pontil's brand

Claude tool calling lets the model decide when to invoke functions you define, then pass structured arguments back for your code to run. By the end of this guide you'll have a working tool-calling loop against the Anthropic Messages API, know how to design schemas the model will actually use, handle the pause-and-resume dance correctly, run tools in parallel, and avoid the failure modes that catch most teams the first time.

Prerequisites: an Anthropic API key, Python 3.10+ or Node 18+, and a rough idea of the workflow you want the agent to complete. Time required: about 45 minutes to get the loop working end to end, longer to harden it.

Step 1 — Install the SDK and make a plain message call first

Before you add tools, prove the transport works. A tool call is just a message with extra fields — if the base Messages API isn't returning cleanly, tools will make it worse, not clearer.

Install the SDK:

pip install anthropic

Send a bare message:

from anthropic import Anthropic

client = Anthropic()

response = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Say hello."}],
)
print(response.content[0].text)

Expected output: a short greeting. If you get an auth error, fix that now — every tool call inherits the same auth path.

Step 2 — Define one tool with a tight JSON schema

Start with one tool, not five. The model picks tools based on the name, the description, and the input_schema. All three matter. A vague description or a loose schema is the single most common reason Claude picks the wrong tool or fills arguments wrong. There's more depth on this in our guide to writing tool descriptions for LLM agents.

tools = [
    {
        "name": "get_order_status",
        "description": "Look up the current status of a customer order by its order ID. Returns status, last update timestamp, and expected delivery date. Use this when the user asks about the state of a specific order.",
        "input_schema": {
            "type": "object",
            "properties": {
                "order_id": {
                    "type": "string",
                    "description": "The order identifier, e.g. 'ORD-12345'."
                }
            },
            "required": ["order_id"]
        }
    }
]

Rules of thumb: describe when to use the tool, not just what it does. Mark every required field. Constrain string enums where possible. Keep parameter counts under 5 for the first version — you can widen later.

Step 3 — Send the tools with your message and inspect the stop reason

Pass the tool list in the request. Claude will either respond directly or return a tool_use block with the arguments it wants you to run.

response = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=1024,
    tools=tools,
    messages=[
        {"role": "user", "content": "What's the status of order ORD-12345?"}
    ],
)

print(response.stop_reason)
print(response.content)

Expected output: stop_reason will be tool_use, and response.content will contain a ToolUseBlock with name="get_order_status", an id, and input={"order_id": "ORD-12345"}. That id matters — you'll echo it back in the next step.

If stop_reason is end_turn instead, the model chose not to call the tool. That usually means the description didn't make the trigger obvious. Rewrite the description before blaming the model.

Step 4 — Execute the tool and return the result in a user message

The pattern that trips people up: tool results are sent as a user message, not an assistant message. Your code runs the function, then you append the full assistant response and a new user message containing a tool_result block that references the original tool_use_id.

def get_order_status(order_id: str) -> dict:
    # your real lookup here
    return {
        "status": "shipped",
        "last_update": "2026-01-14T09:22:00Z",
        "expected_delivery": "2026-01-16"
    }

tool_use = next(b for b in response.content if b.type == "tool_use")
result = get_order_status(**tool_use.input)

messages = [
    {"role": "user", "content": "What's the status of order ORD-12345?"},
    {"role": "assistant", "content": response.content},
    {
        "role": "user",
        "content": [
            {
                "type": "tool_result",
                "tool_use_id": tool_use.id,
                "content": str(result)
            }
        ]
    }
]

final = client.messages.create(
    model="claude-sonnet-4-5",
    max_tokens=1024,
    tools=tools,
    messages=messages,
)
print(final.content[0].text)

Expected output: a natural-language reply summarising the order status. If you get a 400 error, the tool_use_id is almost certainly missing or mismatched.

Step 5 — Wrap the exchange in a loop for multi-turn tool use

One tool call is rarely enough. Claude may need to chain two or three calls to answer. Wrap the send-execute-append cycle in a loop that terminates when stop_reason is no longer tool_use.

def run_agent(user_message: str, max_iterations: int = 10):
    messages = [{"role": "user", "content": user_message}]

    for _ in range(max_iterations):
        response = client.messages.create(
            model="claude-sonnet-4-5",
            max_tokens=1024,
            tools=tools,
            messages=messages,
        )
        messages.append({"role": "assistant", "content": response.content})

        if response.stop_reason != "tool_use":
            return response

        tool_results = []
        for block in response.content:
            if block.type == "tool_use":
                try:
                    result = TOOL_REGISTRY[block.name](**block.input)
                    tool_results.append({
                        "type": "tool_result",
                        "tool_use_id": block.id,
                        "content": str(result)
                    })
                except Exception as e:
                    tool_results.append({
                        "type": "tool_result",
                        "tool_use_id": block.id,
                        "content": f"Error: {e}",
                        "is_error": True
                    })

        messages.append({"role": "user", "content": tool_results})

    raise RuntimeError("Max iterations exceeded.")

Always cap iterations. A runaway loop with no ceiling is how tool-calling bills get memorable.

Step 6 — Handle parallel tool calls without breaking things

Claude will often return multiple tool_use blocks in a single response when the calls are independent — for example, looking up three orders at once. Your loop already handles this shape (iterate every tool_use block, collect every tool_result), but how you execute them matters.

Two rules:

  1. Only run calls in parallel if they're side-effect-free or independently safe. A pair of read-only lookups is fine. Two writes to the same resource is not.
  2. Return every tool_result in the same follow-up user message, in a single content array. Splitting them across messages will 400.

We've written a longer piece on parallel tool calls covering safety classification and concurrency caps if you're pushing this hard.

Step 7 — Add error handling, retries, and observability

Production tool calling breaks in three places: the tool itself throws, the model produces invalid arguments, or the API call times out. Handle each explicitly.

  • Tool exceptions: return a tool_result with is_error: true and a short human-readable message. Claude will usually retry with different arguments or ask the user for clarification. See our guide on API error responses for AI agents for what makes an error string recoverable.
  • Schema violations: validate block.input against your JSON schema before executing. Reject with a structured error rather than crashing the loop.
  • API failures: retry with exponential backoff on 429 and 5xx. Cap retries at 3.

Log every tool_use block, every result, every stop_reason, and total token counts per turn. When something goes wrong in week three, the trace is what saves you.

Common pitfalls

  • Missing tool_use_id on the result. Every tool_use must be matched by exactly one tool_result with the same id. If you skip one, the next request 400s.
  • Returning tool results as assistant role. They're user. Every time.
  • Vague tool descriptions. "Gets order info" tells the model nothing about when to call it. Say what triggers it and what it returns.
  • Loose schemas. "type": "string" with no enum, no format, no pattern will get you garbage inputs. Constrain what you can.
  • No iteration cap. A tool that keeps returning "try again" errors will loop until the context window explodes. Cap it.
  • Treating tool calling as function calling on autopilot. Claude's implementation differs from OpenAI's in shape and semantics. If you're comparing, our post on tool calling vs function calling covers where the contracts diverge.
  • Skipping evals. Tool selection accuracy degrades silently as you add tools. Grade it deliberately, not by vibes.

Join our weekly newsletter

Stay up to date on the ever changing agentic landscape.

POSTS

Related content

Agent infrastructure

Agents in production

OpenAI tool calling: a step-by-step guide to shipping it in production

8 minute read

Agent infrastructure

Agents in production

Tool calling vs function calling: the same mechanism, two production realities

8 minute read

Agent infrastructure

Platform integration

Tool schema design for AI agents: what actually makes a schema the model can use

9 minute read