Agents in production
Agent infrastructure
Tool calling reliability breaks across five layers — selection, arguments, auth, execution, and interpretation. Here's what actually fails and how to fix it.

Tool calling looks solved in the demo and stalls in production. The model picks a tool, fills the arguments, the API returns a result, and the loop closes. Then it goes to real traffic and the failure rate climbs — sometimes to a quarter of calls, sometimes higher — and nobody can agree whether the model, the schema, or the backend is at fault.
Our view: tool calling reliability is not a model problem. It's a system problem with five layers, and each layer has its own failure mode. Treating the whole thing as "the LLM got it wrong" is why teams keep tuning prompts when the fix is somewhere else entirely.
This deep-dive walks through the five layers where tool calls actually break — selection, argument construction, transport and auth, execution, and result interpretation — then covers how to measure reliability honestly and what a production-grade recovery loop looks like. If your agent works in eval and misbehaves under load, one of these layers is where the answer lives.
Before any fix, you need to know which layer is failing. A tool call passes through five distinct stages, and each stage has failure modes that look nothing alike in the logs. Confusing them is the single most common reason agent teams over-invest in prompt engineering and under-invest in the tools layer.
The rest of this article works layer by layer. The order matters: fixing execution reliability while selection is broken means you're making the wrong calls faster.
Selection failures are the most visible and the most misdiagnosed. The model calls create_invoice when the user asked to send a reminder. Or it calls nothing at all and answers from parametric knowledge. Teams see this and reach for a better prompt. Sometimes that works. Often it doesn't, because the problem is upstream.
The two structural causes are catalogue size and description quality. Past roughly 30–50 tools loaded into context, Anthropic has published guidance that selection accuracy degrades, and OpenAI recommends keeping tool counts per turn well under their 128-function ceiling — the practical observation is that accuracy drops faster than most teams expect. Semantic retrieval helps, but only if the retriever's embeddings actually match how the model reasons about the task. Tool descriptions that describe what the tool is rather than when to call it compound the problem: the model has no discriminator to work with when two tools sound similar.
The fix is upstream of the prompt. Curate the tool surface: expose the workflow, not the schema. If you have 200 API endpoints, you probably need 20 tools. Write descriptions in the imperative — "Use this when the user wants X" — and add negative guidance where two tools overlap. We've written more on agent tool selection failures and how many tools an agent should actually have — the short version is that granularity and description quality do more for selection accuracy than any prompt change.
Don't grade on final answer. Grade on whether the correct tool was called, in the right order, with the right frequency. tool_call_precision (of the tools the model called, how many were the right one) and tool_call_recall (of the tools it should have called, how many did it) are the two numbers that matter. If precision is high and recall is low, your descriptions are too narrow. If recall is high and precision is low, they're too broad.
The model picks the right tool and fills the arguments wrong. This is where a lot of "the LLM is unreliable" reports actually live, and it's the failure mode most sensitive to schema design.
JSON Schema is the contract between the model and the tool. What lives in that schema — required vs optional fields, enums vs free strings, nested objects vs flat parameters — decides whether the model can construct valid arguments at all. Three patterns break argument construction more than anything else:
status: "active" | "paused" | "cancelled" gives the model something to pick from; status: string gives it something to hallucinate.customer_id with no way to look one up first will fail on every call where the ID isn't already in context. Either pair the tool with a discovery tool, or accept a natural-language identifier and resolve it server-side.Strict mode helps. OpenAI's Responses API supports strict: true, which enforces the schema at generation time rather than validation time. Anthropic's tool use supports similar constrained decoding. Both meaningfully reduce malformed arguments — but neither fixes a schema that was ambiguous to begin with. See our guide on tool schema design for the shape that actually holds.
The model made a valid call. The tool layer serialised it correctly. The request left the runtime. And then a 401 came back.
Auth is where a surprising share of "tool calling reliability" incidents actually resolve. The patterns are consistent:
The fix is at the runtime, not the model. Tool calls should execute as the authenticated user with per-call, scoped credentials — not a shared account. Refresh should be automatic, not something the tool code has to think about. When permission fails, the runtime should surface a structured error ("missing scope: invoices:write") that the model can actually reason about, not a generic 403. We covered the mechanics in AI agent authentication and least privilege for AI agents.
The request landed. Now the reliability problem shifts from the model to the backend, and the failure modes look like normal distributed-systems problems — with one twist. The caller is an agent, which means retry behaviour is stochastic and can produce duplicate side effects at rates human clients never generate.
The execution-layer failure modes worth naming:
Timeouts. The API is slow, the runtime cuts the call at 30s, the model retries, the original write eventually lands. Now there are two invoices. Idempotency keys are the fix, and they have to be generated by the runtime, not the model — models are terrible at generating stable keys across retries.
Partial writes. The tool writes to two systems and one fails. Without a compensating action, state is now inconsistent and the model has no way to know. Either wrap multi-system tool calls in a durable execution primitive, or expose a check_state tool the model can call to reconcile.
Rate limits. 429s from downstream APIs are the most predictable failure and still catch teams out. Per-user rate limits (from your own auth model) and per-tenant rate limits (from the downstream API) compound — an agent burst can exhaust both simultaneously. Backoff has to be applied at the runtime, not the model loop, or you burn tokens waiting.
Schema drift. The API silently changed the shape of its response. The tool still returns 200 OK. The model reads a field that no longer exists and either fabricates a value or fails downstream. This is the failure mode nobody catches until a customer complains. We covered it in OpenAPI spec drift — the short version is that runtime contract checks are the only reliable defence.
Everything above describes what breaks between the model and the API. Pontil operates on that layer.
Pontil is a Tools-as-a-Service platform. We generate tools from the APIs a product already has, run them through a managed runtime that handles auth, retry, idempotency, and observability, and keep the tool contracts current as the underlying product changes. Tool calls execute as the authenticated user, so permissions and audit trails honour the real identity — not a shared service account. When the API contract drifts, the runtime catches it before the model does.
The reliability layers in this article — selection, arguments, auth, execution, result interpretation — are the ones the tools layer is supposed to hold. If your team is spending its time on prompt tuning when the failures are actually at the boundary, the layer beneath the model is where the leverage is.
The last layer is the one teams instrument least. The API returned. Now the model has to read the response, decide whether the call succeeded, and act on the result.
Three failure modes dominate. False success: the API returned 200 with a body that indicates the operation didn't happen ("queued," "pending human approval"), and the model reads it as done. Truncated context: the response is 40KB, the tool returns all of it, and the model misses the relevant field somewhere in the middle. Infinite retry: the response contains an error the model doesn't recognise, so it calls the same tool again with slightly different arguments — five, ten, twenty times before something breaks the loop.
The fixes are all in the response shape. Return structured, machine-first status fields — status: "completed" | "queued" | "failed" — not English prose the model has to parse. Cap response size at the tool layer and paginate — see API pagination for AI agents. Use structured error contracts with a taxonomy the model can reason about, following the pattern in API error responses for AI agents. Include remediation hints in error responses ("retry after 30s", "missing scope", "no such record") so the model has something to act on instead of guessing.
Most "reliability" numbers agent teams cite are hollow. "95% success rate" usually means 95% of calls returned 2xx, which measures the transport layer and nothing else. Real reliability has to grade the full loop.
The metrics that actually matter, at each layer:
Instrument all five. Report all five separately. A single blended "reliability score" hides the layer where the fix lives, which is the layer teams need to see. Agent tracing platforms are getting better at this — see AI agent observability — but the discipline of separating the layers has to come from the team, not the tool.
Reliability isn't a property of the model or the API in isolation. It's a property of the layer between them — the tools layer — where selection, arguments, auth, execution, and interpretation all have to hold at once. The teams that ship agents to production are the ones that stop asking "why did the model do that?" and start asking "which of the five layers just failed?"
The reason this article isn't a checklist is that the fixes at each layer look nothing alike. Selection is a curation problem. Arguments are a schema problem. Auth is an identity problem. Execution is a distributed-systems problem. Interpretation is a response-shape problem. A team optimising one layer can't move the others, and a team measuring one blended score can't tell which one is broken.
The good news is that the layers are addressable. None of them require a foundation model breakthrough. All of them require somebody to own the tools layer as a real piece of infrastructure — with contracts, tests, observability, and the boring engineering discipline that any production system needs. The agent projects that ship are the ones that treated it that way from the start.
Stay up to date on the ever changing agentic landscape.