Agent infrastructure

Agents in production

AI agent architecture diagram: what actually belongs at each layer in production

AI agent architecture diagram deep-dive: the five layers of a production agent stack, where the standard diagram misleads, and what belongs in the tools layer.

10 minute read
Decorative imagery showcasing Pontil's brand

Most AI agent architecture diagrams you see in slide decks are wrong. Not wrong in the details — wrong in the layers. They draw the model at the top, an orchestrator underneath, a box labelled "tools" or "integrations," and something vague at the bottom about APIs or data. That diagram is what got drawn before anyone tried to ship an agent against a real product.

This piece takes the question seriously: what does a production AI agent architecture diagram actually look like, and what belongs at each layer? The short version of our view: the reasoning layer isn't the interesting part any more, and the tools layer is where most agent projects quietly die. The diagram most teams draw hides that.

The sections below walk through the five layers that matter, the boundaries between them, where the diagram usually gets simplified into uselessness, and what a diagram looks like when it survives contact with production.

Why most agent architecture diagrams mislead

The standard enterprise agent architecture diagram is a stack of five or six boxes. Model. Orchestrator. Memory. Tools. Data. APIs. Arrows going up and down. It's the same diagram everyone draws, and it's misleading in a specific way: it treats every layer as equal in surface area and equal in difficulty. In production, they aren't close.

The reasoning layer — the model — is a commodity. You can swap Anthropic for OpenAI for Google in an afternoon, and most production agents will behave similarly enough that the choice comes down to cost and latency, not capability. The orchestration layer is real work, but frameworks like LangChain, LangGraph, and CrewAI have converged on similar primitives. That layer is solved enough to buy.

The tools layer isn't. And the tools layer is where agents actually do things. That mismatch — a small box in the diagram, most of the engineering effort in reality — is why so many diagrams give a false picture of where the risk lives. We wrote about this at length in the orchestrator obsession is hiding the real bottleneck, and it applies just as much to how the architecture gets drawn.

The rest of this piece proposes a different way to draw it — one where the box sizes match the engineering reality.

The five layers of a production agent architecture

Here's the layer breakdown we've found holds up under real deployments. Each row is a layer, top to bottom, in the order data flows on a typical turn.
‍

Layer
What it does
What it doesn't do

1

Interface

Receives the user request, streams the response, holds the session

Decide what tools exist, execute anything

2

Reasoning

Picks the next action: tool call, response, or clarification

Guarantee the tool call is safe, retryable, or authorised

3

Orchestration

Runs the agent loop, tracks state, handles multi-step plans, retries

Know how any specific tool actually behaves

4

Tools

Exposes product capabilities as invocable functions, executes calls under the user's identity, handles auth and errors

Generate the plan or reason about the outcome

5

Systems of record

The actual product, database, or API being acted on

Care whether the caller is an agent or a human


Every production agent has these five layers, whether the team drew them or not. The interesting question is which ones you build, which you buy, and which one determines whether the project ships at all.

Interface

The interface is the thinnest layer. It's a chat UI, a Slack app, an in-product agent panel, an API endpoint another service calls. It handles session state, streaming, and rendering. Nothing about the interface is agent-specific — the same patterns hold whether you're building a chatbot or a form.

Teams that overinvest here usually don't have a real agent problem yet. If your reasoning layer is picking the wrong tools, no amount of interface polish fixes that.

Reasoning

The reasoning layer is the model plus the prompt plus the tool definitions the model sees on this turn. It decides: given what the user asked and what tools are available, what should happen next? A tool call, a clarifying question, or a final response.

In 2024 this layer was interesting because model capability was the ceiling. In 2026 it isn't, mostly. Modern frontier models pick tools well when the tool descriptions are good and pick them badly when they aren't. That's a tools-layer problem, not a reasoning problem. We covered the failure modes in agent tool selection: why the model picks the wrong tool.

Orchestration

Orchestration is the loop. It takes the model's decision, dispatches the tool call, feeds the result back, and asks the model what to do next. It tracks conversation history, handles retries at the loop level, and manages branching for multi-agent or multi-step workflows.

This is where LangGraph, OpenAI's Agents SDK, CrewAI, and Anthropic's agent patterns live. They're differentiated by ergonomics, not by fundamental capability. Most teams should pick one and move on — the orchestrator isn't where the project succeeds or fails.

Where it does matter: durable execution for long-running workflows, human-in-the-loop checkpoints, and multi-agent coordination. If any of those describe your agent, orchestration is real engineering. If your agent is a single-turn tool caller, orchestration is a library import.

Tools

The tools layer is what turns a model's decision into an action against a real system. It has three responsibilities that the diagram usually collapses into one box:

  1. Definitions. What tools exist, what their schemas look like, and how the model discovers them.
  2. Runtime. How a tool call actually executes: auth, rate limiting, error handling, observability, idempotency.
  3. Maintenance. How tool definitions stay in sync as the underlying products change.

The reason most agent architecture diagrams break is that they draw "tools" as a single arrow pointing at "APIs." In practice, the tools layer is a full subsystem — often bigger than everything above it combined. It's where the auth model gets designed, where per-user identity flows through, where tool descriptions get written and rewritten, and where the connectors to the underlying products get built and maintained.

This is the layer we spend all of Pontil's engineering effort on, which is why we care about drawing it correctly.

Systems of record

The bottom layer is what everyone else calls "the product" or "the API." It's the CRM, the billing system, the ticketing platform, the internal service. From the agent's perspective, it's the thing that has state and that the agent is trying to read from or write to.

This layer doesn't know or care that an agent is calling it. It exposes whatever surface it exposes — usually an API that was designed for a UI, not for agents. The gap between what this layer can do and what its API allows is the structural problem we wrote about in your APIs expose 2% of what your product can do. The diagram has to show this layer honestly, because pretending the API is the product is where architecture decisions go wrong.

The boundaries between layers matter more than the layers themselves

A useful architecture diagram spends as much time on the lines between boxes as the boxes. Three boundaries are where production agent projects break:

Reasoning ↔ Orchestration. The model returns a tool call; the orchestrator has to trust it enough to dispatch, but not so much that a hallucinated tool name or malformed argument takes down the loop. Every mature orchestration layer has a validation and coercion step here. Skipping it is where you get 3am pages.

Orchestration ↔ Tools. The orchestrator hands off a tool call and gets back a result. What's in that result decides whether the next turn is useful. If the tool returns a 500-line JSON blob, the next turn's reasoning drowns. If it returns an opaque error, the model retries the same call forever. Response shape is a boundary contract, not a tools-layer implementation detail. We covered the specifics in API pagination for AI agents: why cursor beats offset when the caller isn't human.

Tools ↔ Systems of record. This is the boundary that fails silently. The tool calls the underlying API. If that API was designed for a UI — which it was — then critical capabilities aren't reachable, per-user auth doesn't flow through cleanly, and the audit trail collapses into a service-account log. The team building the agent usually doesn't own the underlying API, which means the fix isn't in the agent codebase at all.

Drawing these boundaries with the same visual weight as the layer boxes forces the conversation about who owns what — which is the conversation that determines whether the project ships.

Where the standard diagram breaks in production

Four patterns show up when a whiteboard diagram meets a real deployment.

The tools box is too small. The diagram allocates one row to "tools" and three rows to reasoning, memory, and orchestration. The engineering effort is inverted: the tools row is where 70% of the work is. Redraw it with box sizes proportional to time-to-build and the diagram tells you where to staff.

Auth is drawn as an arrow, not a layer. A production agent has to execute as the authenticated user, not as a shared service account, or the audit trail is worthless and permissions collapse. That's not a line on the diagram — it's a subsystem inside the tools runtime, with token exchange, delegated identity, and per-call authorisation. We went deep on this in AI agent authorization: the boundary that decides whether your agents can ship.

Maintenance is missing entirely. The standard diagram is a snapshot of runtime behaviour. It doesn't show how tool definitions stay accurate as the underlying products change. Every SaaS product ships breaking changes to internal endpoints continuously. If your diagram doesn't show a maintenance loop — CI checks against the upstream API, drift detection, automatic tool regeneration — you've drawn a system that works for a week.

Systems of record are drawn as a monolith. In an established enterprise, there are dozens of them, half-owned, with inconsistent auth models and API coverage. The diagram usually shows one box labelled "APIs." In reality it's a portfolio, and the tools layer has to abstract that portfolio into a coherent surface for the agent — otherwise the agent has a different execution model for every backend, which the reasoning layer can't handle.

What a production-ready diagram actually looks like

A diagram that survives production has five properties:

  1. Layer sizing matches engineering effort. Tools is the biggest layer. Reasoning is a thin bar with a model name in it.
  2. Auth flows are drawn as a vertical band, not a horizontal arrow. Identity has to be present at every layer — the interface authenticates the user, the orchestrator carries the identity through, the tools runtime executes calls under it, and the systems of record enforce permissions on it. If identity drops out at any layer, the whole thing fails security review.
  3. The maintenance loop is explicit. There's an arrow from systems of record back up to tool definitions, showing that tool schemas track the products they're generated from. Without it, the diagram is a photograph of a moment, not a system.
  4. Boundaries have contracts, not just arrows. Each boundary between layers is labelled with what the contract is: JSON tool call format, tool response shape, delegated auth token type, upstream API version. When a boundary breaks in production, the label tells you who owns the fix.
  5. The tools layer has three sub-boxes: definitions, runtime, maintenance. Collapsing them into one hides the biggest source of risk in the whole system.

A diagram with those five properties is longer and messier than the pretty stack in the slide deck. It's also the one that predicts where the project will stall.

How Pontil fits

When we talk to teams whose agent projects have stalled, the diagnosis is almost always in the tools layer — and specifically in the three sub-boxes we called out above. Definitions don't exist for the capabilities the agent needs. The runtime doesn't carry per-user identity end to end. Maintenance is manual, so tool schemas rot within weeks of shipping.

Pontil is a Tools-as-a-Service platform, and that's the layer we operate on. We generate tool definitions from the codebases and APIs a product already has, run them in a managed runtime that executes as the authenticated user, and keep them current as the underlying products change. It's not a framework, an orchestrator, or an integration platform — it's the tools layer of the diagram, built as a product so it doesn't have to be built from scratch by every team.

If you want the longer version of why so many agent projects stall at this exact boundary, we wrote it up in why agent projects stall.

What changes when you draw the diagram honestly?

The standard agent architecture diagram is a communication tool for people who haven't tried to ship an agent yet. It's fine for that. It stops being fine the moment your team has to decide where to staff, what to build versus buy, and where the compliance risk sits.

A production-ready diagram makes three things visible that the standard version hides. It shows that the tools layer is where most of the engineering lives. It shows that identity and maintenance are structural properties of the whole stack, not features of one box. And it shows that systems of record aren't a monolith — they're a portfolio the tools layer has to make coherent.

Drawing it that way changes the conversation. The question stops being "which framework should we pick" and starts being "who owns the tools layer, and how do we keep it honest as the products underneath us change." That's the question worth spending Q3 on.

Join our weekly newsletter

Stay up to date on the ever changing agentic landscape.

POSTS

Related content

Agent infrastructure

Platform integration

Orchestrator vs tools layer: where agent work actually happens

7 min read

Agent infrastructure

The agent stack: a map for platform teams

6 minute read

Agents in production

Platform integration

Agent readiness framework: how to measure whether your SaaS is ready for AI agents

7 minute read