Platform integration
Agents in production
Bidirectional sync infinite loop prevention: how to add origin tagging, filter echo events, resolve conflicts, and add a runtime loop breaker in 7 steps.

Bidirectional sync between two systems sounds simple until the first write echoes back and triggers itself. You update a contact in your product, the change fires a webhook to the CRM, the CRM writes it, the CRM fires a webhook back to you, you write it again, and the loop is running. By the end of this guide you'll know how to detect an infinite loop, add origin tagging to every write, filter echo events at the ingestion boundary, and handle the edge cases where tagging alone isn't enough. It assumes you already have a working bidirectional sync between two systems and a webhook receiver on each side. Time required: 60–90 minutes to implement, longer to test properly.
Before fixing anything, prove the loop exists. A lot of "loops" are actually double writes, not infinite recursion. Both are bad. Only one is a loop.
Run this check. Turn on structured logging for both sync directions. Pick one record. Write to it once from System A. Then count events over the next 60 seconds.
If you see the second pattern, you have a loop. If you see the first, you have an idempotency problem — different fix, covered in webhook idempotency key: a practical guide to exactly-once processing. This guide assumes the second.
Origin tagging is the foundation of loop prevention. Every write to either system must carry a marker that says "this write came from the sync, not from a user." When you receive a webhook, you check the marker. If it says the write came from you, you drop the event.
There are three places to put the tag, ranked by reliability:
sync_origin: "pontil-sync" on every write.metadata object on many of its updateable objects (Customer, PaymentIntent, Subscription, etc.), and that metadata is included in the resulting webhook event payloads.actor.id, and you filter on it. Least clean but works everywhere.Pick the highest option your target system supports. If none work, you're falling back to Step 4 (content hashing).
Here's the tagging pattern using a service account:
# In your sync writer
response = crm_client.update_contact(
contact_id=contact_id,
fields=payload,
auth=SYNC_SERVICE_ACCOUNT_TOKEN, # not the user's token
)
On the receiving side, the very first thing your webhook handler does — before parsing, before enqueueing, before anything — is check the origin tag. If the tag says the event was caused by your own sync, you drop it.
def handle_webhook(event):
# Echo filter runs first. Before signature verification's business logic.
# Before anything that costs money or writes state.
if is_echo(event):
log.info("dropped_echo", event_id=event["id"])
return 200 # ACK so the sender doesn't retry
# Now proceed with the normal pipeline
return process(event)
def is_echo(event):
# Depending on which tagging strategy you picked in Step 2:
if event.get("actor", {}).get("id") == SYNC_SERVICE_ACCOUNT_ID:
return True
if event.get("metadata", {}).get("sync_origin") == "pontil-sync":
return True
if event.get("data", {}).get("sync_origin") == "pontil-sync":
return True
return False
Two things worth calling out. Return 200, not 400 — the sender did nothing wrong, and returning an error triggers retries, which look like a loop. And log every dropped echo, because when the filter breaks you want to know immediately.
Some systems strip custom fields and metadata before firing webhooks. Some don't propagate the actor identity in the payload. When origin tagging doesn't survive the round trip, you need a content-based fallback: hash the payload you just wrote, store the hash with a short TTL, and drop incoming events whose payload hashes match.
import hashlib
import json
def payload_hash(record_id, fields):
canonical = json.dumps({"id": record_id, "fields": fields}, sort_keys=True)
return hashlib.sha256(canonical.encode()).hexdigest()
# On write: store the hash with a 5-minute TTL
redis.setex(f"recent_write:{payload_hash(id, fields)}", 300, "1")
# On receive: check if the incoming payload matches something we just wrote
def is_echo_by_content(event):
incoming_hash = payload_hash(event["record_id"], event["fields"])
return redis.exists(f"recent_write:{incoming_hash}")
Content hashing is a fallback, not a first choice. Two writes with identical content from different sources will get dropped as false positives. Keep the TTL short — long enough to cover the delivery latency you actually see, and no longer.
Echo filtering stops the loop. It doesn't stop the harder problem underneath: two systems editing the same field at the same time. If a user edits a contact's phone number in System A while a different user edits it in System B, both writes fire webhooks. You've filtered your own echoes, so both non-echo writes land on the opposite system. Now each side has stale data and fires again. That looks like a loop, but it's a conflict.
The fix is to pick a conflict resolution strategy per field, not per record:
Store the last-modified timestamp with every synced field, not per record. When an incoming event tries to update a field, compare its timestamp against the local one and apply the strategy. Connector sync conflict resolution happens here, not in the echo filter.
Echo filtering and conflict resolution should stop every loop you can imagine. Add a runtime loop breaker anyway. It's the seatbelt.
Count the number of sync events per record per rolling window. If any single record fires more than a threshold — say, 10 sync events in 5 minutes — halt sync for that record and page someone.
def check_loop_breaker(record_id):
key = f"sync_count:{record_id}"
count = redis.incr(key)
if count == 1:
redis.expire(key, 300)
if count > 10:
halt_sync(record_id)
alert(f"Sync loop breaker tripped for {record_id}")
return False
return True
The breaker doesn't fix loops — it stops them from destroying your rate limits, your database, and your reputation with the downstream API while you're asleep. Related infrastructure patterns: the circuit breaker pattern for APIs.
Write integration tests that specifically try to create loops. Manual verification is not enough — echo filters break silently when a target system changes its webhook payload format.
At minimum, test:
Run these on every deploy. When they break, they break silently in production, and you find out from a customer with a wrecked CRM.
Returning 4xx or 5xx for echo events. The sender treats an error as "try again," which produces a retry storm that looks exactly like a loop. Always return 200 for a dropped echo.
Filtering by user identity instead of service account. If your sync writes as the user who triggered the change, echoes are indistinguishable from that user's own next edit. Sync must write as its own identity.
Assuming webhook order matches write order. Most webhook providers do not guarantee ordering. Two rapid writes from your sync can arrive as webhooks in reverse order. Origin filtering doesn't care, but any logic that reasons about "which write came last" needs the timestamp from the payload, not the receive time.
Skipping the loop breaker because "filtering works." Filtering works until a target system changes its payload format and drops your tag field. The loop breaker is what saves you when the filter silently stops matching.
Testing only in one direction. Loops are bidirectional by definition. Every test above needs to run A→B and B→A.
Stay up to date on the ever changing agentic landscape.