Error handling & retries
As events flow through Chord they are processed (enrichment, transformation, filtering) and then delivered to one or more destinations. Either stage can fail. This document explains how Chord classifies those failures, which ones are retried, how retries are scheduled, and what happens when a connection's property mappings are misconfigured.
Delivery Model
Chord provides at-least-once delivery. When a recoverable failure occurs, the affected event is re-queued and delivered again later. Because of this, an event may occasionally be delivered more than once — see Deduplication for how destinations should handle duplicates using messageId.
Each event is delivered independently per connection (a source → destination pairing). A failure delivering an event to one destination does not affect delivery of the same event to any other destination.
Recoverable vs Unrecoverable Errors
Chord divides failures into two categories, and only one of them is retried.
Recoverable (transient) errors are failures that are expected to resolve on their own without any change to configuration or code — for example a destination API returning a 5xx, a request timing out, a rate-limit response, or a temporary network blip. These errors are retried.
Unrecoverable errors are failures that retrying cannot fix — for example malformed event data a destination will never accept, an authentication failure, or a logic error in a custom function. These are not retried. Depending on the failure the event is dropped, recorded as failed, or (for a function that throws a standard error) still delivered — and in each case a notification is sent to any configured connection monitors.
The table below summarizes how the most common outcomes are handled.
Outcome | Source | Retried | Result |
|---|---|---|---|
Transient delivery failure (5xx, timeout, rate limit, network error) | Destination | Yes | Re-queued with backoff; delivered once the destination recovers |
Recoverable failure raised by a function | UDF (RetryError) | Yes | Re-queued with backoff |
Permanent destination rejection (4xx, auth/validation failure) | Destination | No | Not delivered to that destination, monitors notified |
Unrecoverable failure raised by a function | UDF (NoRetryError) | No | Event dropped, monitors notified |
Unexpected logic error | UDF (standard Error) | No | Event still delivered, monitors notified |
Intentional filtering | UDF (return null / false / []) | No | Event silently dropped, no notification |
Functions (UDFs) can deliberately raise either a retryable or a non-retryable error to control this behavior. That mechanism is documented in detail in Functions; this document covers only how the pipeline acts on those signals.
How Retries are Scheduled
When an event fails with a recoverable error, Chord does not retry it immediately. Instead it computes a future retry time and re-queues the message on a dedicated retry topic. The message waits there until its scheduled retry time passes, at which point it is reprocessed.
Retries use exponential backoff so that a struggling destination is given progressively more time to recover between attempts:
- A fixed number of retry attempts is made per event (default: 3).
- The delay before each attempt grows exponentially from a base interval (default base: 10 minutes), so successive attempts are delayed roughly 10 minutes, then ~1.7 hours, then ~16.7 hours.
- Every delay is capped at a maximum (default: 24 hours), so no single retry is ever scheduled further out than the cap.
These values are operational defaults and may be tuned per environment.
On a retry, processing resumes at the stage that failed rather than from the beginning. A failure during destination delivery re-attempts delivery only, without re-running the function pipeline. A failure inside the function pipeline re-runs the whole function pipeline together with destination delivery. This avoids duplicating the side effects of stages that already completed successfully.
The Dead Letter Queue
Retries are not infinite. Once an event has exhausted its retry attempts and still cannot be delivered, it is moved to a dead-letter queue rather than being retried forever or silently discarded. This keeps the live pipeline healthy while preserving the failed event for inspection and potential manual replay.
A move to the dead-letter queue is recorded in the pipeline metrics and surfaced to any configured connection monitors, so persistent delivery failures are visible rather than hidden.
Property Mapping Problems
A misconfigured connection property mapping does not stop the event. Mappings are applied per connection, so a bad mapping affects only that one destination — every other destination still receives the event as normal.
A source path that does not exist on the event is the most common misconfiguration: you map properties.loyalty_tier but the event never carries it, or it only carries it on some events. When that happens, that single mapping is skipped, a warning is logged, and the event is delivered with all of its other mappings applied. The destination simply does not receive that one field.
Where the warning appears depends on how the destination runs:
- Browser (device-mode) destinations — pixels and tags that run in the visitor's browser — log a warning to the browser console. It names the destination and the mapping, and ends with source path "<path>" does not exist on event — mapping skipped. These warnings do not appear in Live Events.
- Server-side (cloud-mode) destinations log the same warning to Live Events in the Chord Console.
A warning on every event usually means the mapping is too broad rather than wrong. Use the mapping's Event Scope to restrict it to the event types and names that actually carry the property, and the warnings stop.
Static-value mappings never warn, because there is no source property to look up.
In the rare case that a mapping fails in an unexpected way rather than simply missing its source, the error is logged and that event is not sent to that destination. Other destinations are unaffected. If you see this, contact Chord support — it indicates a problem on our side, not in your configuration.
Monitoring Failures
Recoverable and unrecoverable errors alike are recorded in the pipeline's metrics and, where connection monitors are configured, generate notifications. You can inspect individual failures and the errors that caused them in Live Events within the Chord Console.
Intentional filtering (an event dropped by a function returning a falsy value) is recorded as dropped and does not generate a notification — it is treated as normal pipeline behavior, not a failure.
Related Topics
- Functions — how custom functions raise retryable and non-retryable errors
- Deduplication — handling the duplicate deliveries that at-least-once retries can produce