Event deduplication
Delivery Guarantee
The Chord CDP provides at-least-once delivery for every event. This means the CDP attempts to deliver each event at least once and may retry when failures occur, so in rare cases the same event may be delivered more than once.
Duplicates are typically caused by transient events in the underlying message pipeline:
- A processing instance restarting (deployment, autoscaling, crash) before its progress is acknowledged
- A network blip or pause that causes an in-flight message to be redelivered
- An upstream producer retry after an unconfirmed delivery
- Temporary failures in downstream services that trigger replays
This is the same delivery model used by the major customer data platforms (Segment, Rudderstack, mParticle), and is the standard guarantee for streaming event pipelines.
Current Deduplication Policy
The Chord CDP does not perform message-level deduplication within the pipeline itself. Instead, deduplication is the responsibility of the downstream destination.
For every event the CDP processes, the source-provided messageId is preserved end-to-end and made available to every destination. Destinations should use the messageId (or a deterministic value derived from it — some destinations require a transformed identifier, e.g., a messageId with a suffix) to detect and reject duplicates on their side.
How destinations should handle duplicates
Destination Type | Recommended Approach |
|---|---|
Data warehouses (Snowflake, BigQuery, Postgres, Redshift, ClickHouse, etc.) | Configure deduplication at the destination using messageId as the primary key. The Chord CDP supports per-connection deduplication options (deduplicate + primaryKey) that perform MERGE/UPSERT on each load. ClickHouse destinations collapse duplicates on background merges via the ReplacingMergeTree table engine. |
HTTP APIs (Braze, Klaviyo, Insider, Stripe, etc.) | Pass event.messageId (or a destination-specific derivation of it) as an idempotency key in the API request — typically via the Idempotency-Key HTTP header. Most modern APIs will reject or ignore requests with a previously-seen idempotency key. |
Reverse ETL / loopback destinations | Filter by messageId in the destination system, or rely on the destination's natural primary-key constraints. |
Why Deduplicate at the Destination?
Destination-side deduplication is the conventional pattern for streaming data systems, for several reasons:
- Destinations have authoritative state. A warehouse already knows whether a row with a given primary key exists. An HTTP API already knows whether it has processed a given idempotency key. Asking these systems to detect duplicates is more reliable than maintaining a parallel record elsewhere.
- Destinations are diverse. Different destinations have different definitions of "duplicate" — some merge by primary key, some upsert by composite key, some collapse on background processes. A pipeline-level dedup can't capture this nuance.
- It avoids a single point of failure. A pipeline-level dedup store would be a critical-path dependency; if it became slow or unavailable, the entire pipeline would degrade. Pushing dedup to destinations keeps the pipeline fast and stateless.
- It aligns with how streaming pipelines work. At-least-once is the default delivery semantic of the underlying streaming infrastructure. Building exactly-once on top requires complex coordination that introduces its own failure modes.
Implications for Custom Functions (UDFs)
UDFs run inside the CDP pipeline before events are dispatched to destinations. Because the CDP is at-least-once, a UDF may execute more than once for the same source event in rare redelivery scenarios.
For UDFs that are pure transformations (enrichment, filtering, splitting), this is harmless — the destination will still deduplicate the result by messageId.
However, UDFs that perform external write operations (calling a third-party API that modifies state, incrementing a remote counter, sending an email, writing to a database) should be designed to be idempotent. The standard approach is to pass event.messageId as an idempotency key to the external system:
export default async function (event, { fetch }) {
await fetch("https://api.example.com/orders", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Idempotency-Key": event.messageId, // ← prevents duplicate writes on redelivery
},
body: JSON.stringify({ /* ... */ }),
});
return event;
}Without an idempotency key, a redelivered event could cause the UDF's external write to occur twice.
Working with messageId
Every event flowing through the CDP carries a messageId field. If the source provides one, it is preserved end-to-end. If not, the CDP generates a unique identifier at ingest time. The messageId is:
- Stable — the same value is forwarded to every destination
- Unique — within a reasonable window (system-generated unique identifiers, or source-provided identifiers that the source guarantees unique)
- Available everywhere — exposed on the event object inside UDFs as event.messageId, included in the payloads sent to destinations, and visible in Live Events
Use it as the canonical key for any deduplication, tracing, or correlation downstream.
Related Topics
- For details on configuring per-connection deduplication options for a warehouse destination, see your destination configuration page in the Chord Console.
- For guidance on writing idempotent UDFs, see the CDP Functions documentation.
Support
For questions about deduplication, delivery guarantees, or how to configure your destination, please contact [email protected].