Agents in production

Agent infrastructure

AI agent memory governance: when to remember, when to forget, and who decides

AI agent memory governance decides whether agents survive production. Why drift is the default, when agents should forget, and where policy actually lives.

5 minute read
Decorative imagery showcasing Pontil's brand

Most agent memory conversations start in the wrong place. They start with what the agent should remember. The harder question — the one that decides whether the agent survives a security review or a support ticket six months in — is what it should forget, and who gets to decide.

We think memory governance is the next boundary agent teams will have to hold. Not memory as a feature. Memory as a policy surface, with the same weight as auth and audit.

Memory is being treated like a product feature. It shouldn't be

The framing in most agent tutorials is that memory makes agents smarter. Store the user's preferences, the past conversations, the intermediate results, and the next interaction gets better. That framing is fine for a demo. It falls apart the moment the agent is running against real user data inside a multi-tenant SaaS product.

At that point memory stops being a feature and starts being a liability. Every fact the agent remembers is a fact someone can subpoena, a fact that has to survive a GDPR deletion request, a fact that can leak across sessions, and a fact that ages badly. A preference from March is a bug in October. A cached permission from before a role change is a security incident.

The agent memory vs RAG vs tools comparison already argues that most production agents need tools before they need memory. The follow-up is: if you do add memory, treat it as governed state, not as a smarter cache.

Drift is not a bug, it is the default

Memory drift is what happens when what the agent remembers stops matching what is true. It is not an edge case. It is the resting state of any memory system that outlives the fact it stored.

Three kinds of drift show up in production:

  • Stale facts. The user changed their default project, their team, their timezone, their manager. The agent still acts on the old one.
  • Stale permissions. The user's role was downgraded. The agent's memory of what they could do wasn't.
  • Stale context. The workflow the agent remembers optimising for was deprecated. The agent keeps optimising for it.

Every one of these looks fine in evals and breaks quietly in production. The agent doesn't error. It just does the wrong thing confidently, which is worse. And because memory is usually written once and read many times, one bad write compounds across every future session until something explicit clears it.

The teams we've spoken with who take this seriously have stopped treating memory as write-and-forget. They treat every memory entry as a fact with a shelf life, an owner, and a source of truth it must reconcile against.

Agent memory policy needs three things most teams don't have

A working agent memory policy is not a paragraph in a design doc. It is three concrete decisions, made per memory type, enforced at runtime.

Retention. How long does this memory live before it is re-verified or dropped? Session-scoped memory dies at session end. User-preference memory needs a TTL and a re-confirmation path. Permission-adjacent memory should probably not be memory at all — it should be a fresh lookup, every call.

Provenance. Where did this fact come from? User told the agent directly, agent inferred it from behaviour, agent read it from a system of record. The three have different trust levels and different rules for when they can be overwritten.

Deletion. Who can delete this, and what triggers automatic deletion? A GDPR request has to reach every active memory store, not just the primary database (backups can be handled via the 'put beyond use' route rather than immediate deletion, but active stores are not optional). A role change has to invalidate cached authorisation state. An account close has to wipe agent memory as thoroughly as it wipes the account.

Most teams have one of the three. Very few have all three, wired to the same enforcement point. The gap is where incidents live.

When agents should forget

The interesting design question isn't when to remember. It's when to forget on purpose.

We think the default should be aggressive forgetting, with narrow exceptions for memory that has earned its keep. That inverts the current default, which is store-everything-and-hope. Aggressive forgetting looks like:

  • Forget between sessions unless there is an explicit reason to persist. Most agent context does not need to survive the conversation. Session state is not memory.
  • Forget when the source of truth changes. If the CRM was updated, the agent's cached view of the CRM is wrong. Invalidate it. Don't refresh it lazily.
  • Forget when the user's identity or permissions change. Not "refresh eventually." Forget, and re-derive from the current authoritative state.
  • Forget when the memory hasn't been used. A fact that hasn't been read in ninety days is a fact the agent doesn't need. It is also a fact that is almost certainly stale.
  • Forget on request, everywhere. One deletion API, one policy, every memory store honours it.

The teams that get this right treat forgetting as a first-class capability, with the same operational rigour as writes. The teams that don't end up with agents that carry around six months of decayed context and act on it.

The runtime is where governance actually lives

Memory policy written in a design doc is not memory governance. Memory governance is what the runtime enforces on every read and every write — the TTL check, the provenance stamp, the identity binding, the deletion propagation. If the enforcement lives in application code, it will be inconsistent across agents. If it lives in the runtime that serves every tool call and memory access, it holds.

This is the same pattern that shows up in AI agent authorization and in least privilege for AI agents. The boundary that decides whether the system holds is the runtime one. Policy without a runtime is a promise. Policy with a runtime is a control.

Agent teams tend to build memory before they build the runtime that governs it. That order is backwards. Decide who enforces retention, provenance, and deletion — then decide what to remember.

If you take one thing from this piece: memory drift is not something you fix once. It is a permanent operating condition of any agent that outlives a single session. The teams that ship agents that don't embarrass them treat forgetting as a product decision, not a cleanup task. Start there, and the rest of memory design gets easier.

Join our weekly newsletter

Stay up to date on the ever changing agentic landscape.

POSTS

Related content

Agents in production

Agent infrastructure

Agent memory vs RAG vs tools: which one your agent actually needs

8 minute read

Agent infrastructure

Platform integration

Least privilege for AI agents: what the boundary actually looks like in production

10 minute read

Agent infrastructure

Platform integration

AI agent authorization: the boundary that decides whether your agents can ship

10 minute read