Agent infrastructure
Agents in production
Tool chaining is how agents sequence dependent tool calls. What actually breaks in production, why the model isn't the real problem, and how to design past it.

Tool chaining is the mechanism that turns a single-tool demo into an agent that can actually finish work. It's also where most production agent projects break — quietly, in ways that look like model failures but aren't.
The question this article answers: what does it take to make sequential tool calls hold up when the second call depends on the first, the third depends on the second, and any one of them can return something the model didn't expect? Pontil's view in one sentence: tool chaining is a runtime discipline, not a prompting technique — the calls only compose reliably when the tools themselves are designed to hand off state cleanly, and when the runtime around them handles the failure modes the model can't reason about.
This piece walks through five things: what tool chaining actually is (and how it differs from parallel calls), why the model isn't your real problem, the four failure modes that kill chains in production, how to design tools that chain well, and what the runtime layer has to do underneath.
Tool chaining is the pattern where an agent makes a sequence of tool calls in which each call depends on the output of a previous one. The agent calls search_customers to find a customer ID, feeds that ID into get_customer_orders, picks an order from the result, then calls refund_order with that order's ID. Three calls, three dependencies, one workflow.
This is not the same as parallel tool calls, where the model fires several independent calls in one turn because none of them need each other's output. Parallel calls are a concurrency problem. Chained calls are a state problem. The model has to hold intermediate results in its context, pick the right field from the previous response, and pass it into the next call without hallucinating the value or dropping the type.
The terminology varies. Some teams call this sequential tool calls, some call it dependent tool calls, some call it tool call chaining in the context of an LLM specifically. The OpenAI Responses API documents it under multi-turn tool use. Anthropic's Claude API documents the same pattern under tool_use blocks and tool_result blocks alternating across turns. MCP doesn't standardise chaining at all — it standardises the transport, and leaves the orchestration to the client. So when a blog post says "tool chaining," it usually means one of two things: the sequence of tool calls a single agent makes across turns to complete a task, or a pre-defined pipeline where tool B is wired to consume tool A's output regardless of what the model decides. This article is about the first one. The second is closer to a workflow engine, and the trade-offs there — durability, retries, compensation — are worth their own piece (see our durable execution deep-dive for that).
The distinction matters because most teams reach for tool chaining when they realise a single tool can't do the whole job. That's the right instinct. The wrong instinct is assuming the model will figure out the sequencing if the tools are individually correct. It won't.
The temptation, when a chain breaks, is to blame the model. It picked the wrong field. It hallucinated an ID. It forgot the context. It called the tools in the wrong order.
Sometimes that's true. Modern models — Claude 4, GPT-5, Gemini 2.5 — chain three or four dependent calls reliably when the tools are well-designed. Push past five or six and error rates climb regardless of the model. But in the projects we see, the model is rarely the first thing that broke. The first thing that broke was one of these:
Each of these is a tool design or response shape problem. The model didn't fail — it faithfully executed on inputs that were structurally ambiguous. Which is why swapping in a smarter model rarely fixes a broken chain. It just changes which call fails first.
This is the piece most teams miss. The reliability of a tool chain is bounded by the weakest handoff in it, and the handoffs are properties of the tools, not the model.
Across the agent projects we've dug into, chain failures cluster into four patterns.
The most common failure. Tool A returns an object with an id field. Tool B takes a customer_id parameter. The model, most of the time, correctly maps id to customer_id. Sometimes it passes email instead. Sometimes it passes a display name. Sometimes it passes the ID from the wrong nested object in a list response.
The fix isn't better prompting. It's shaping tool responses so the identifier the next tool needs is unambiguous. Return customer_id in tool A's response, not id. Use consistent field names across the whole tool surface. If a tool returns multiple objects, structure the response so the model doesn't have to guess which one is the primary result.
Search returns zero results. Get returns null because the record was deleted between calls. List returns exactly one item where the model expected several. Each of these is a valid response the tool must handle explicitly.
The worst failure mode is the tool returning an empty response with a 200 status and no explanation. The model has no signal that anything went wrong, so it continues the chain with garbage inputs. The fix: return structured empty responses with a status or reason field the model can read. {"status": "no_match", "reason": "No customer found with email X"} gives the model something to reason about. [] gives it nothing.
The API returns 200 but the operation only partially completed. A bulk update processed 8 of 10 records. A payment authorised but didn't capture. A workflow started but hit a downstream validation error the API surfaced in a warnings field the model didn't parse.
This is a design problem in the underlying API, not just the tool wrapping it. But at the tool layer, you can shape the response to hoist warnings into a place the model will read. Don't bury warnings at the end of a large nested object — put it at the top of the response, and prefix descriptions with explicit language like "partial success" that the model is likely to attend to.
The model assumes the state observed in tool A still holds when tool C runs. It usually does. Sometimes it doesn't — because another user made a change, because a webhook fired, because time passed. The chain proceeds on stale assumptions and produces a wrong result.
This is where idempotency and preconditions matter. Destructive tools should accept an if_version or expected_state parameter and refuse to execute if the underlying state has moved. The model doesn't have to reason about concurrency; the tool refuses the operation and returns a structured conflict response the model can react to.
Notice that only the last row involves the runtime at all. The other three are tool design problems that runtime can't paper over.
The goal is tools whose outputs are the inputs of other tools without the model having to translate. Three principles.
Consistent identifiers, consistently named. If your customer object has an ID, call it customer_id everywhere it appears — in the search response, in the get response, as the parameter name in every tool that takes a customer as input. The model's job is to pick the field with the matching name. Don't make it do type inference.
Return the shape the next tool expects. If search_customers returns a list and the natural next call is get_customer, the search response should return customer_id values in a form get_customer can accept directly. If list_orders is expected to feed into refund_order, return order_id values, not URLs or composite keys that need parsing. The tool schema is the contract; make it a contract the model can honour without inference.
Make status legible. Every tool response should include some structured indication of what happened. Success, no results, partial success, refused-due-to-precondition. The model reads response text, so a status field with a small controlled vocabulary is more useful than a boolean success flag hidden in a nested object.
The underlying pattern here is HATEOAS applied to agent tools — the response tells the caller what to do next. In a hypermedia API that means links. In an agent tool that means field names and status values that map cleanly onto the next call in the chain.
One more thing: keep tools small enough that chains have to exist. There's a temptation, once you've watched a chain break, to fuse the three tools into one super-tool that does the whole workflow. That works for the specific workflow you fused. It breaks the moment the workflow needs a variation the fused tool doesn't support. Tool granularity is a trade-off worth thinking about explicitly, not a problem to route around by inlining logic into ever-larger tools.
Even with well-designed tools, chains fail. The runtime — the layer that actually executes tool calls between the model and the underlying APIs — has to handle the failures the model can't.
Retry with backoff on transient errors. A 503 from an upstream API shouldn't propagate to the model as "the tool failed." The runtime should retry with exponential backoff, only surfacing the failure to the model after retries are exhausted. When it does surface, it should surface as a structured error the model can react to, not a stack trace.
Enforce timeouts. A tool that hangs blocks the whole chain. The runtime needs per-tool timeout limits, and when a timeout fires it should return a structured timeout response, not silence.
Preserve identity across the chain. If the chain executes as an authenticated user, every call in the chain must execute as that user. Not as a shared service account, not as the identity of whichever token happened to be cached. This matters more the deeper the chain goes, because a chain that starts with search_my_customers and ends with refund_order is only correct if the entire chain honours the same permission scope. See our piece on zero trust for agents for the identity-boundary argument in full.
Log the chain, not just the calls. When a chain fails in production, the useful debugging artefact isn't a single failed tool call — it's the whole trace, with inputs, outputs, and the model's reasoning between calls. Agent observability that groups tool calls into chains is the difference between a debuggable system and a mystery.
Guard against runaway chains. The model can loop. It calls tool A, gets a result it doesn't like, calls tool A again with a small variation, still doesn't like it, calls tool A again. Without a cap on chain depth or a mechanism to detect repeated calls with near-identical inputs, the runtime will happily execute forty tool calls that accomplish nothing. Set a chain depth limit. Detect loops. Break with a structured error the model can respond to.
None of this is glamorous. All of it is what separates a demo from a production system.
Tool chaining is where the two halves of Pontil's product surface — generated tools and the runtime that executes them — stop being separate concerns. Chains only compose reliably when the tools are shaped to hand off state cleanly, and when the runtime around them handles retries, timeouts, identity, and loop detection consistently across every call in the sequence.
Tools-as-a-Service is our answer to both sides at once. Tools are generated from your existing codebase with consistent identifier naming, structured status responses, and precondition support built into the contract. The runtime executes them as the authenticated user, retries and times out on standard rules, and traces the whole chain — not just the individual calls — into the observability layer. That combination is what makes a five-step chain reliable enough to ship, rather than a demo that works twice and fails the third time.
The honest answer is: it depends on which of the four failure modes above you haven't addressed yet. In our experience, identifier drift and empty-response handling break first, silent partial success breaks second, and state assumptions break last — but only because most teams never get their chains long enough to hit the state problem. Fix the first two and you can usually chain four or five calls reliably. Fix all four and the ceiling moves to eight or ten, at which point context window management becomes the next bottleneck.
What doesn't change is the shape of the work. Tool chaining is a runtime discipline. Better prompts don't fix broken handoffs. Better models don't fix ambiguous responses. The chain gets more reliable when the tools get better designed and the runtime gets stricter about what it will and won't paper over. That's true for two-step chains and it's more true for ten-step ones.
Worth watching over the next year: whether protocols like MCP start to standardise anything about chaining explicitly, whether models start exposing structured chain planning rather than just next-turn tool selection, and whether the runtime layer around agents starts to look less like a proxy and more like a workflow engine. Our bet is on the third. The other two would help, but the work of making dependent tool calls hold up in production is going to sit at the runtime layer either way.
Stay up to date on the ever changing agentic landscape.