Agents in production

Agent infrastructure

Shadow mode AI agent deployment: what it catches, what it hides, and how to run it well

Shadow mode AI agent deployment done right: what it catches, the four failure modes that mislead, and how to design a shadow harness that produces real signal.

9 minute read
Decorative imagery showcasing Pontil's brand

Shadow mode is the deployment pattern most teams reach for when their agent works in eval but nobody trusts it in production. The agent runs against real traffic, produces real decisions, and touches nothing. You compare what it would have done to what actually happened, catch the divergences, and iterate. Then, when the numbers look right, you flip the switch.

That's the pitch. It's a good pitch. But shadow mode for agents behaves differently from shadow mode for ML models, and most of the frameworks teams inherit from the ML side quietly break when the thing being shadowed is a tool-calling agent instead of a classifier.

Our view: shadow mode is worth doing, but only if you're honest about what it can't measure. This piece covers what shadow mode is when the caller is an agent, the four failure modes that make naïve shadow deployments misleading, how to design the comparison layer so it produces signal instead of noise, and what to do when shadow mode says green but production says red.

What shadow mode actually is when the caller is an agent

Shadow mode — sometimes called shadow deployment, dark launch, or prediction-only mode — runs a new system in parallel with the live one. The new system sees the same inputs. It produces outputs. Those outputs are logged and compared against the live system's outputs, but never acted on.

For an ML model, this is straightforward. A fraud classifier sees a transaction, emits a score, and you compare that score to the incumbent's. No side effects. No state. The shadow model is a pure function of its input.

Agents are not pure functions. An agent that would have called create_invoice in shadow mode didn't create an invoice, which means the next step in its trajectory — the one that reads the invoice ID back — never runs. Or runs against stale data. Or hallucinates the ID. The shadow trajectory diverges from the live trajectory the moment the first tool call is skipped, and everything after that point is measuring something that wouldn't have happened.

This is the first thing to be clear about. Shadow mode for an agent isn't "what would the agent have done?" It's "what would the agent have done up to the first side-effect it wanted to have?" After that, you're running a simulation, and the fidelity of that simulation depends entirely on how well you can fake the world.

The teams that get shadow mode right treat this as the design constraint, not a footnote. The teams that get it wrong ship a shadow harness, watch the metrics, and get blindsided when the real launch surfaces failures the shadow never had a chance to see.

Four failure modes that make naïve shadow deployments misleading

The standard shadow-mode setup for agents looks like this: mirror the inbound request, run the new agent alongside the incumbent (or alongside a human), log every tool call the agent wanted to make, compare. Four things go wrong.

Trajectory drift after the first skipped side-effect

Covered above, but worth restating in operational terms. If your agent's trajectory is search → read → decide → write → confirm, and you're shadowing the write, the confirm step is now reading a world where the write didn't happen. The agent will notice. It will either re-attempt the write (double logging), give up (false negative), or hallucinate success (silent corruption of your comparison data).

Every shadow run past the first suppressed side-effect is measuring a counterfactual, not the agent. Treat trajectory length as a confidence interval — the further past the first suppressed call you look, the less the comparison means.

Read-write asymmetry in the tool surface

Most real agents mix reads and writes. Shadow mode wants to allow the reads (so the agent has real context) and block the writes (so nothing changes). This sounds simple. It isn't.

Some "reads" have side effects — GET endpoints that log analytics events, list operations that touch rate-limit counters, retrievals that warm caches. Some "writes" are idempotent enough to be safe if you can guarantee the same request never fires twice. And the boundary between "read" and "write" isn't in the HTTP method; it's in what the downstream system does with the call. A tool schema that doesn't distinguish safe reads from unsafe reads makes correct shadowing impossible.

The comparison layer becomes the bottleneck

If your shadow harness compares agent output to a human's output, you now need a human in every loop, producing labels at the rate the agent produces trajectories. That's fine at 10 requests a day and unworkable at 10,000. If it compares to the incumbent system, you need the incumbent to be running the same input against the same context — which usually means blocking on it, which means you've turned an async shadow into a synchronous latency doubling.

The comparison layer is where most shadow deployments quietly stop being real. Teams start sampling. Then they sample less. Then the labels stop keeping up. Then the shadow is just a log of what the agent wanted to do, with nobody looking at it.

Auth context that doesn't match production

If the shadow agent runs under a different identity than the production caller — a service account instead of the real user's delegated token — then everything auth-adjacent is untested. Permission errors won't fire. Row-level filters won't apply. Rate limits will hit different quotas. The shadow reports "the agent would have succeeded," and production reports "the agent got a 403," and both are correct about their own world.

This is why agent identity has to match user identity all the way through the shadow path. If the runtime can't execute shadow calls as the authenticated user, the shadow is measuring the wrong thing.

How to design a shadow deployment that produces signal

Given those failure modes, the useful shadow-mode design has five properties.

Classify every tool by side-effect class, not HTTP method. For each tool, decide: safe to execute in shadow (pure read, no analytics, no rate-limit impact), safe to simulate (deterministic response you can generate from a fixture), or unsafe to shadow at all (execution changes state you can't reason about). Tools in the third category are trajectory-ending — the shadow run stops there and you record what the agent wanted to do without pretending you know what would have happened next.

Run the shadow agent under the real user's delegated credentials. Same OAuth token, same scopes, same tenant context as the live call. The shadow does everything the live agent would do at the auth layer, up to the point where a write would fire. This is the only way permission errors, tenant boundaries, and quota behaviour show up honestly. If your runtime can't do this cleanly, that's the problem to fix before the shadow is worth running.

Log the full trajectory, not just the outcome. Every tool considered, every tool selected, every argument constructed, every response received (or simulated), every reasoning step between them. When shadow mode says the agent picked the wrong tool, you need the trace to figure out why. Outcome-only logs tell you the diff without telling you the cause — which is exactly the debugging position you're trying to avoid. This is where agent evals grade trajectories, not just final outputs.

Sample deliberately, weight by risk. You cannot compare every shadow trajectory to ground truth. Sample high-value flows at 100%, long-tail flows at whatever your labelling budget allows, and accept that the shadow is a spot-check on the tail, not an audit. Weight your confidence intervals accordingly when you present numbers to the people making the launch decision.

Set a trajectory-length cutoff. Past N tool calls after the first suppressed side-effect, stop comparing. The signal is gone. Report divergence rate for trajectories that didn't hit that cutoff, and separately report how often trajectories hit it — because a shadow that keeps hitting the cutoff isn't measuring the agent's behaviour, it's measuring your inability to shadow it.

The metrics that actually matter

What people track
What actually predicts production

Tool selection

Match rate against incumbent

Match rate on high-value tools, weighted by cost of a wrong pick

Argument construction

Exact-match on JSON

Semantic match on required fields; ignore ordering and defaults

Trajectory

Full-path match

Divergence point + reason for divergence

Latency

Shadow p95

Shadow p95 *plus* the cost of the simulated-response layer, which won't exist in production

Auth failures

Rarely tracked

4xx rate under real user tokens, per tenant


The left column is what shadow dashboards ship with. The right column is what tells you whether the agent will hold when the flag flips.

How Pontil fits

Most of what makes shadow mode hard for agents is a tools-layer problem. Whether a tool is safe to execute in shadow depends on whether its side-effect surface is honestly described. Whether the shadow runs under the real user's identity depends on whether the runtime can execute tool calls as the authenticated user rather than a shared service account. Whether trajectories are debuggable depends on whether every tool call is traced with enough fidelity to reconstruct what the agent saw.

This is the layer Pontil operates in. Tools-as-a-Service means the connectors an agent invokes are generated from the systems you own, executed at runtime under the caller's delegated identity, and instrumented at the tool-call boundary. The side-effect classification, the auth flow, and the trace granularity that a well-run shadow deployment needs are the same properties a well-run production deployment needs — because shadow mode is production, minus the writes you chose to suppress.

If your shadow harness is producing noise instead of signal, the fix is rarely in the harness. It's in the tools layer underneath it.

What happens when shadow says green and production says red

Even a well-designed shadow deployment will occasionally clear an agent that then misbehaves in production. When it happens, the failure almost always sits in one of three places.

The agent's real trajectory involves a tool interaction the shadow never observed, because the shadow's simulated responses masked the class of failure that shows up when a real system returns something unexpected. Fix: widen the fixture set for simulated tools, or reclassify borderline tools as trajectory-ending instead of simulated.

The agent's live behaviour depends on state the shadow didn't have access to. A common version: the shadow ran under the same user's token but at a different point in time, and the world had moved. Fix: shrink the shadow-to-live gap, and be more sceptical of comparisons across long-running trajectories.

The agent's failure mode is a low-frequency, high-cost event that shadow sampling missed. This is the hardest one, because it's not a shadow-design failure — it's a statistics failure. Fix: don't rely on shadow mode alone for tail-risk decisions. Pair it with a staged rollout that gives you the same visibility on live traffic without the same blast radius.

The honest posture on shadow mode is that it de-risks the middle of the distribution and tells you very little about the tail. Treat it as a filter, not a proof.

Where does shadow mode fit in the broader agent rollout playbook?

Shadow mode is one tool among several: pre-deployment eval suites, shadow deployment, canary rollouts with fractional traffic, feature-flagged launches per tenant, and full rollout with continuous eval. Each stage catches a different failure class, and none of them catch everything.

The pattern we see in teams shipping agents to production reliably isn't a linear pipeline. They run evals continuously, keep a shadow lane running in perpetuity for regression detection, canary every model or tool update, and accept that "production" is a state you keep verifying rather than one you reach. The pattern isn't launch-then-monitor. It's launch-into-a-monitoring-system that never stops.

Shadow mode earns its place in that pattern when it's honest about its limits. Used as a first-and-only gate, it will mislead you. Used as one signal among several, feeding a rollout process that assumes production traffic will surface things no shadow could, it's one of the higher-leverage patterns available. The question isn't whether to run it. It's whether the tools layer underneath is honest enough to make what it measures worth acting on.

Join our weekly newsletter

Stay up to date on the ever changing agentic landscape.

POSTS

Related content

Agent infrastructure

Agents in production

AI agent error handling: a practical guide to retries, circuit breakers, and recovery

8 minute read

Agents in production

Agent infrastructure

Agent evals: how to measure tool calls, trajectories, and production reality

9 minute read

Agent infrastructure

Agents in production

Durable execution for AI agents: what it is, when you need it, and where it breaks

9 minute read