Agents in production
Platform integration
Agentic AI coding tools split into IDE-native and terminal-based. How each approach works, where each wins, and how to pick the one that fits your team.

A senior engineer on a team we work with lost half a day last month to a Claude Code run that looked fine until it wasn't. The task was routine: reconcile a schema drift between a staging Postgres and the ORM models. The agent had MCP access to the staging DB, read/write on the repo, and a plan that read cleanly in the terminal. Twenty minutes in, a migration step failed. The agent retried. The retry hit a different code path in the MCP server, which quietly re-ran an earlier ALTER TABLE against a table it had already modified. By the time anyone looked, the staging schema had a column duplicated under a slightly different name and three test fixtures were referencing the wrong one. Recovery took a rollback, a manual diff, and a conversation about why the audit log showed the same tool call three times with no correlation ID tying them to a single plan step.
Nothing about that incident was exotic. It's the failure mode you get when an autonomous agent has real tool access, a retry loop, and no control plane sitting between the plan and the side effects. And it's the reason the split between IDE-native coding assistants and terminal-based coding agents is worth paying attention to — not as a buying decision for your engineering team, but as a live case study in the two architectures you're choosing between every time you ship an agent into your own product.
On one side sit the IDE-native assistants — Cursor, GitHub Copilot, Windsurf — that live inside the editor and act on the surface a human is already looking at. On the other sit the terminal-based coding agents — Claude Code, Aider, OpenAI Codex CLI — that run as autonomous processes, reach further into your systems, and expect to be handed a task rather than a keystroke.
Both call themselves agentic. The interesting question for a Head of AI or Head of Product is not which one your devs should use. It's what each pattern implies about auth, tool exposure, and runtime when you build the same shape of agent into your own product.
IDE-native agents are the reference implementation of an in-app agent surface. Cursor forks VS Code and rebuilds the sidebar around a chat panel. Copilot ships as an extension. Windsurf takes a similar shape with Cascade. In all three, the agent inherits the user's session, the user's file system permissions, and the user's view of the world. The model reads what the user can already see: open buffers, workspace, LSP, diagnostics.
The interaction pattern is turn-by-turn. The user describes a change, the agent proposes a diff, the user accepts or rejects. Destructive actions get approved one at a time. The human is in the loop at the granularity of a single edit or command.
Translate that to a product agent — a Copilot-style panel inside your SaaS — and the properties carry over. Auth is whatever the user is already logged in as. Tool exposure is bounded by what the current view can already do. The audit story is easier because every side effect is tied to an explicit user gesture. The failure mode is scope: the agent is shaped around one user and one context, and it struggles when the task spans multiple services or needs to reach systems the user's session doesn't already touch.
This is the pattern most product teams reach for first, and for good reason. It's the cheapest to get right on the control-plane axis.
Terminal-based agents are the reference implementation of an autonomous backend agent. You launch Claude Code in a repo, give it a prompt, and it plans, edits, runs tests, and reports back. Aider works the same way. OpenAI's Codex CLI takes a similar shape. The agent has a shell, file system access, and — through MCP servers or plugin systems — access to whatever else you wire in: databases, ticketing systems, CI logs, deployment APIs.
The interaction pattern is task-by-task. You describe a job — "migrate this module from callbacks to async/await and update the tests" — and the agent works through it in a longer loop. Approval granularity is coarser: a plan and a change-set, not each edit.
Translate that to a product agent — a backend worker that triages tickets, reconciles data, or drives a multi-step workflow across your integrations — and the properties change sharply. The agent no longer inherits a user session, so auth becomes a design problem: service accounts, delegated tokens, per-tool scopes. Tool exposure is whatever you chose to wire in, which means the blast radius is whatever the loosest MCP server allows. The audit story has to be built, not inherited. And the failure mode is the one from the postmortem above: silent drift inside a plan the human isn't watching turn by turn.
The in-app pattern is forgiving because the user is the control plane. The autonomous pattern is not, and this is where most product teams underestimate the work.
Give an agent MCP access to a staging DB and four decisions come due at once:
Auth and identity. The agent is not a user. It needs its own principal — a service account, a scoped token, or a delegated identity that borrows a user's permissions for a bounded window. Most teams start with a shared token that has too much access and no expiry, discover the problem the first time an agent writes to the wrong environment, and rebuild. Design the identity model before you wire the first tool.
Per-tool scoping. MCP servers advertise capabilities at the server level, but your permissioning needs to live at the tool level. read_table and alter_table are the same MCP server and radically different risk surfaces. The teams that stay out of trouble treat every tool as a separate grant, defaulted to deny, with a written justification for each one that's on.
Idempotency and retry semantics. The staging-DB incident above happened because a retry re-executed a side effect. Any tool the agent can call more than once needs an idempotency key or a server-side dedupe window. Agents will retry. Loops will re-enter. Assume it and design the tool contract to be safe under repetition — this is one of the failure modes we cover in tool schema design for AI agents.
Audit and correlation. The single most useful thing you can add to an autonomous agent's runtime is a correlation ID that ties every tool call to the plan step that issued it, and every retry to the original call. Without it, the audit log is a flat list of tool calls and you cannot reconstruct why. With it, a postmortem takes an hour instead of a day. Log the plan, log the step, log the call, log the result — with one ID threading through all of them.
The teams that get stuck are almost always the ones that treated these as ops concerns to solve later. They are product concerns, and they shape what the agent can and can't be trusted to do.
Reach for the in-app agent surface when:
Reach for the autonomous backend agent when:
Most mature products end up with both, wired to the same tool layer underneath. The in-app surface handles the interactive work. The backend agent handles the long-running or cross-system work. What makes that combination tractable is a shared connectivity layer that both agents call into, with the auth and audit story solved once instead of twice.
The useful thing about watching Cursor and Claude Code sit side by side is that they make the architectural choice legible. One inherits a user's context and pays for that with limited reach. The other reaches everywhere and pays for it with a control-plane problem the user can't solve for you.
When you ship an agent into your own product, you are making the same trade. And whichever side you pick, the ceiling is set by the same thing: what your agent can safely reach, under what identity, with what audit trail, through what tool contracts. The model matters less than most teams think. The connectivity layer matters more.
That's the work we spend most of our time on with product teams — the integration backlog and the MCP surface that sits under the agent, not the agent itself. If you're mid-decision on this, our writing on MCP server setup and architecture and how agents pick between tools is the closest thing to the shape of that work. The agent pattern you ship is downstream of it.
Stay up to date on the ever changing agentic landscape.