Platform integration

Agents in production

Bidirectional sync infinite loop: how to stop echo events from ping-ponging between systems

Bidirectional sync infinite loop prevention: how to add origin tagging, filter echo events, resolve conflicts, and add a runtime loop breaker in 7 steps.

8 minute read
Decorative imagery showcasing Pontil's brand

Bidirectional sync between two systems sounds simple until the first write echoes back and triggers itself. You update a contact in your product, the change fires a webhook to the CRM, the CRM writes it, the CRM fires a webhook back to you, you write it again, and the loop is running. By the end of this guide you'll know how to detect an infinite loop, add origin tagging to every write, filter echo events at the ingestion boundary, and handle the edge cases where tagging alone isn't enough. It assumes you already have a working bidirectional sync between two systems and a webhook receiver on each side. Time required: 60–90 minutes to implement, longer to test properly.

Step 1 — Confirm you actually have a loop

Before fixing anything, prove the loop exists. A lot of "loops" are actually double writes, not infinite recursion. Both are bad. Only one is a loop.

Run this check. Turn on structured logging for both sync directions. Pick one record. Write to it once from System A. Then count events over the next 60 seconds.
‍

Pattern
What you're seeing
What it is

A→B, B→A, then silence

Two events, one per direction

Double write, not a loop

A→B, B→A, A→B, B→A…

Events keep firing on the same record

Infinite loop

A→B, then B→A five minutes later

Delayed second event

Probably a scheduled reconciliation job, not a loop


If you see the second pattern, you have a loop. If you see the first, you have an idempotency problem — different fix, covered in webhook idempotency key: a practical guide to exactly-once processing. This guide assumes the second.

Step 2 — Tag the origin of every write

Origin tagging is the foundation of loop prevention. Every write to either system must carry a marker that says "this write came from the sync, not from a user." When you receive a webhook, you check the marker. If it says the write came from you, you drop the event.

There are three places to put the tag, ranked by reliability:

  1. A dedicated field on the record. Cleanest option if the target system lets you add custom fields. Set sync_origin: "pontil-sync" on every write.
  2. A metadata field on the API request. Some APIs support request-level metadata that flows through to the resulting webhook. Stripe does this with the metadata object on many of its updateable objects (Customer, PaymentIntent, Subscription, etc.), and that metadata is included in the resulting webhook event payloads.
  3. A dedicated service account or API key. Every sync write authenticates as a separate identity. The webhook payload includes actor.id, and you filter on it. Least clean but works everywhere.

Pick the highest option your target system supports. If none work, you're falling back to Step 4 (content hashing).

Here's the tagging pattern using a service account:

# In your sync writer
response = crm_client.update_contact(
   contact_id=contact_id,
   fields=payload,
   auth=SYNC_SERVICE_ACCOUNT_TOKEN,  # not the user's token
)

Step 3 — Filter echo events at the ingestion boundary

On the receiving side, the very first thing your webhook handler does — before parsing, before enqueueing, before anything — is check the origin tag. If the tag says the event was caused by your own sync, you drop it.

def handle_webhook(event):
   # Echo filter runs first. Before signature verification's business logic.
   # Before anything that costs money or writes state.
   if is_echo(event):
       log.info("dropped_echo", event_id=event["id"])
       return 200  # ACK so the sender doesn't retry
   
   # Now proceed with the normal pipeline
   return process(event)

def is_echo(event):
   # Depending on which tagging strategy you picked in Step 2:
   if event.get("actor", {}).get("id") == SYNC_SERVICE_ACCOUNT_ID:
       return True
   if event.get("metadata", {}).get("sync_origin") == "pontil-sync":
       return True
   if event.get("data", {}).get("sync_origin") == "pontil-sync":
       return True
   return False

Two things worth calling out. Return 200, not 400 — the sender did nothing wrong, and returning an error triggers retries, which look like a loop. And log every dropped echo, because when the filter breaks you want to know immediately.

Step 4 — Add content hashing for systems that strip metadata

Some systems strip custom fields and metadata before firing webhooks. Some don't propagate the actor identity in the payload. When origin tagging doesn't survive the round trip, you need a content-based fallback: hash the payload you just wrote, store the hash with a short TTL, and drop incoming events whose payload hashes match.

import hashlib
import json

def payload_hash(record_id, fields):
   canonical = json.dumps({"id": record_id, "fields": fields}, sort_keys=True)
   return hashlib.sha256(canonical.encode()).hexdigest()

# On write: store the hash with a 5-minute TTL
redis.setex(f"recent_write:{payload_hash(id, fields)}", 300, "1")

# On receive: check if the incoming payload matches something we just wrote
def is_echo_by_content(event):
   incoming_hash = payload_hash(event["record_id"], event["fields"])
   return redis.exists(f"recent_write:{incoming_hash}")

Content hashing is a fallback, not a first choice. Two writes with identical content from different sources will get dropped as false positives. Keep the TTL short — long enough to cover the delivery latency you actually see, and no longer.

Step 5 — Handle bidirectional conflicts explicitly

Echo filtering stops the loop. It doesn't stop the harder problem underneath: two systems editing the same field at the same time. If a user edits a contact's phone number in System A while a different user edits it in System B, both writes fire webhooks. You've filtered your own echoes, so both non-echo writes land on the opposite system. Now each side has stale data and fires again. That looks like a loop, but it's a conflict.

The fix is to pick a conflict resolution strategy per field, not per record:
‍

Strategy
When to use it

Last write wins

Low-stakes fields where recency matters (status, notes)

First write wins

Immutable fields where the initial value is canonical (external IDs)

Source of truth

Fields owned by one system by policy (billing amount from the billing system)

Manual merge

High-stakes fields where automatic resolution creates business risk


Store the last-modified timestamp with every synced field, not per record. When an incoming event tries to update a field, compare its timestamp against the local one and apply the strategy. Connector sync conflict resolution happens here, not in the echo filter.

Step 6 — Add a runtime loop breaker

Echo filtering and conflict resolution should stop every loop you can imagine. Add a runtime loop breaker anyway. It's the seatbelt.

Count the number of sync events per record per rolling window. If any single record fires more than a threshold — say, 10 sync events in 5 minutes — halt sync for that record and page someone.

def check_loop_breaker(record_id):
   key = f"sync_count:{record_id}"
   count = redis.incr(key)
   if count == 1:
       redis.expire(key, 300)
   if count > 10:
       halt_sync(record_id)
       alert(f"Sync loop breaker tripped for {record_id}")
       return False
   return True

The breaker doesn't fix loops — it stops them from destroying your rate limits, your database, and your reputation with the downstream API while you're asleep. Related infrastructure patterns: the circuit breaker pattern for APIs.

Step 7 — Test the loop conditions before shipping

Write integration tests that specifically try to create loops. Manual verification is not enough — echo filters break silently when a target system changes its webhook payload format.

At minimum, test:

  1. Sync write does not echo. Write from System A through the sync. Assert System B fires a webhook. Assert System A's handler drops that webhook. Assert no second write happens.
  2. User write does propagate. Write from System A as a real user. Assert System B fires a webhook. Assert System A's handler does not drop it (or does drop it, once, after receiving B's echo). Assert final state is consistent.
  3. Content hash fallback works. Simulate a webhook that has no origin tag but matches a recent payload hash. Assert it's dropped.
  4. Loop breaker trips. Force-fire 15 events for the same record in a minute. Assert sync halts and the alert fires.

Run these on every deploy. When they break, they break silently in production, and you find out from a customer with a wrecked CRM.

Common pitfalls

Returning 4xx or 5xx for echo events. The sender treats an error as "try again," which produces a retry storm that looks exactly like a loop. Always return 200 for a dropped echo.

Filtering by user identity instead of service account. If your sync writes as the user who triggered the change, echoes are indistinguishable from that user's own next edit. Sync must write as its own identity.

Assuming webhook order matches write order. Most webhook providers do not guarantee ordering. Two rapid writes from your sync can arrive as webhooks in reverse order. Origin filtering doesn't care, but any logic that reasons about "which write came last" needs the timestamp from the payload, not the receive time.

Skipping the loop breaker because "filtering works." Filtering works until a target system changes its payload format and drops your tag field. The loop breaker is what saves you when the filter silently stops matching.

Testing only in one direction. Loops are bidirectional by definition. Every test above needs to run A→B and B→A.

Join our weekly newsletter

Stay up to date on the ever changing agentic landscape.

POSTS

Related content

Platform integration

API strategy

Webhook idempotency key: a practical guide to exactly-once processing

7 minute read

Platform integration

Agent infrastructure

Webhook retry with exponential backoff: a practical implementation guide

7 minute read

Agents in production

Platform integration

Queue-first webhook architecture: why the receiver should never do the work

9 minute read