Agents in production

Platform integration

Agentic AI coding tools: how the two dominant approaches actually compare

Agentic AI coding tools split into IDE-native and terminal-based. How each approach works, where each wins, and how to pick the one that fits your team.

8 minute read
Decorative imagery showcasing Pontil's brand

Two agent architectures, and what they teach product builders about shipping agents

A senior engineer on a team we work with lost half a day last month to a Claude Code run that looked fine until it wasn't. The task was routine: reconcile a schema drift between a staging Postgres and the ORM models. The agent had MCP access to the staging DB, read/write on the repo, and a plan that read cleanly in the terminal. Twenty minutes in, a migration step failed. The agent retried. The retry hit a different code path in the MCP server, which quietly re-ran an earlier ALTER TABLE against a table it had already modified. By the time anyone looked, the staging schema had a column duplicated under a slightly different name and three test fixtures were referencing the wrong one. Recovery took a rollback, a manual diff, and a conversation about why the audit log showed the same tool call three times with no correlation ID tying them to a single plan step.

Nothing about that incident was exotic. It's the failure mode you get when an autonomous agent has real tool access, a retry loop, and no control plane sitting between the plan and the side effects. And it's the reason the split between IDE-native coding assistants and terminal-based coding agents is worth paying attention to — not as a buying decision for your engineering team, but as a live case study in the two architectures you're choosing between every time you ship an agent into your own product.

On one side sit the IDE-native assistants — Cursor, GitHub Copilot, Windsurf — that live inside the editor and act on the surface a human is already looking at. On the other sit the terminal-based coding agents — Claude Code, Aider, OpenAI Codex CLI — that run as autonomous processes, reach further into your systems, and expect to be handed a task rather than a keystroke.

Both call themselves agentic. The interesting question for a Head of AI or Head of Product is not which one your devs should use. It's what each pattern implies about auth, tool exposure, and runtime when you build the same shape of agent into your own product.

The in-app agent surface pattern

IDE-native agents are the reference implementation of an in-app agent surface. Cursor forks VS Code and rebuilds the sidebar around a chat panel. Copilot ships as an extension. Windsurf takes a similar shape with Cascade. In all three, the agent inherits the user's session, the user's file system permissions, and the user's view of the world. The model reads what the user can already see: open buffers, workspace, LSP, diagnostics.

The interaction pattern is turn-by-turn. The user describes a change, the agent proposes a diff, the user accepts or rejects. Destructive actions get approved one at a time. The human is in the loop at the granularity of a single edit or command.

Translate that to a product agent — a Copilot-style panel inside your SaaS — and the properties carry over. Auth is whatever the user is already logged in as. Tool exposure is bounded by what the current view can already do. The audit story is easier because every side effect is tied to an explicit user gesture. The failure mode is scope: the agent is shaped around one user and one context, and it struggles when the task spans multiple services or needs to reach systems the user's session doesn't already touch.

This is the pattern most product teams reach for first, and for good reason. It's the cheapest to get right on the control-plane axis.

The autonomous backend agent pattern

Terminal-based agents are the reference implementation of an autonomous backend agent. You launch Claude Code in a repo, give it a prompt, and it plans, edits, runs tests, and reports back. Aider works the same way. OpenAI's Codex CLI takes a similar shape. The agent has a shell, file system access, and — through MCP servers or plugin systems — access to whatever else you wire in: databases, ticketing systems, CI logs, deployment APIs.

The interaction pattern is task-by-task. You describe a job — "migrate this module from callbacks to async/await and update the tests" — and the agent works through it in a longer loop. Approval granularity is coarser: a plan and a change-set, not each edit.

Translate that to a product agent — a backend worker that triages tickets, reconciles data, or drives a multi-step workflow across your integrations — and the properties change sharply. The agent no longer inherits a user session, so auth becomes a design problem: service accounts, delegated tokens, per-tool scopes. Tool exposure is whatever you chose to wire in, which means the blast radius is whatever the loosest MCP server allows. The audit story has to be built, not inherited. And the failure mode is the one from the postmortem above: silent drift inside a plan the human isn't watching turn by turn.

Comparing the two patterns

In-app agent surface
Autonomous backend agent

Primary interaction

Turn-by-turn, user-driven

Task-by-task, agent-driven

Auth model

Inherits the user's session

Service account or delegated token, scoped per tool

Tool exposure

Bounded by the current view's capabilities

Bounded by what you wire into MCP or plugins

Human-in-the-loop granularity

Per-edit / per-action approval

Per-plan / per-change-set approval

Audit trail

Naturally tied to user gestures

Must be constructed; needs correlation IDs across retries

Best-fit task size

Seconds to minutes

Minutes to hours, often unattended

Runtime

User's process / browser tab

Long-lived worker with its own lifecycle

Failure mode

Slow when context is scattered across surfaces

Silent when a plan drifts mid-run

The control plane is where teams get stuck

The in-app pattern is forgiving because the user is the control plane. The autonomous pattern is not, and this is where most product teams underestimate the work.

Give an agent MCP access to a staging DB and four decisions come due at once:

Auth and identity. The agent is not a user. It needs its own principal — a service account, a scoped token, or a delegated identity that borrows a user's permissions for a bounded window. Most teams start with a shared token that has too much access and no expiry, discover the problem the first time an agent writes to the wrong environment, and rebuild. Design the identity model before you wire the first tool.

Per-tool scoping. MCP servers advertise capabilities at the server level, but your permissioning needs to live at the tool level. read_table and alter_table are the same MCP server and radically different risk surfaces. The teams that stay out of trouble treat every tool as a separate grant, defaulted to deny, with a written justification for each one that's on.

Idempotency and retry semantics. The staging-DB incident above happened because a retry re-executed a side effect. Any tool the agent can call more than once needs an idempotency key or a server-side dedupe window. Agents will retry. Loops will re-enter. Assume it and design the tool contract to be safe under repetition — this is one of the failure modes we cover in tool schema design for AI agents.

Audit and correlation. The single most useful thing you can add to an autonomous agent's runtime is a correlation ID that ties every tool call to the plan step that issued it, and every retry to the original call. Without it, the audit log is a flat list of tool calls and you cannot reconstruct why. With it, a postmortem takes an hour instead of a day. Log the plan, log the step, log the call, log the result — with one ID threading through all of them.

The teams that get stuck are almost always the ones that treated these as ops concerns to solve later. They are product concerns, and they shape what the agent can and can't be trusted to do.

When each pattern fits in your product

Reach for the in-app agent surface when:

  • The agent's job maps to something the user is already doing in the UI.
  • Per-action approval is acceptable — often desirable — because the user is present.
  • The auth story is "whatever the user can already do," and you want to keep it that way.
  • The blast radius should be bounded by a single session.

Reach for the autonomous backend agent when:

  • The task is large enough or slow enough that a user won't sit through it.
  • The agent needs to reach systems the user's session doesn't touch — internal APIs, third-party integrations, data stores.
  • You're prepared to build the control plane: identity, per-tool scopes, idempotent tool contracts, correlated audit.
  • The value of running unattended justifies the setup cost. It usually does, once the control plane exists. It rarely does before.

Most mature products end up with both, wired to the same tool layer underneath. The in-app surface handles the interactive work. The backend agent handles the long-running or cross-system work. What makes that combination tractable is a shared connectivity layer that both agents call into, with the auth and audit story solved once instead of twice.

The tool choice is downstream of your connectivity layer

The useful thing about watching Cursor and Claude Code sit side by side is that they make the architectural choice legible. One inherits a user's context and pays for that with limited reach. The other reaches everywhere and pays for it with a control-plane problem the user can't solve for you.

When you ship an agent into your own product, you are making the same trade. And whichever side you pick, the ceiling is set by the same thing: what your agent can safely reach, under what identity, with what audit trail, through what tool contracts. The model matters less than most teams think. The connectivity layer matters more.

That's the work we spend most of our time on with product teams — the integration backlog and the MCP surface that sits under the agent, not the agent itself. If you're mid-decision on this, our writing on MCP server setup and architecture and how agents pick between tools is the closest thing to the shape of that work. The agent pattern you ship is downstream of it.

Join our weekly newsletter

Stay up to date on the ever changing agentic landscape.

POSTS

Related content

Platform integration

Agent infrastructure

Claude Code MCP: how to connect Claude Code to MCP servers

7 minute read

Agents in production

Agent infrastructure

Agent tool selection: why the model picks the wrong tool, and how to design past it

9 minute read

Agent infrastructure

Agents in production

SDK vs API for AI agents: which interface actually holds up in production

7 minute read