Chord Data Source Ingestion Guidelines
Your Data Warehouse → Chord
Getting started with Chord is designed to be frictionless. We meet your data where it is, allowing you to unlock AI-driven insights, leverage the Chord Copilot, and activate your data downstream, all while leveraging your existing infrastructure investments. All we need is for your data to land in cloud storage as files. We take it from there. For Snowflake users, please refer to Sharing Data with Chord via Snowflakehttps://docs.chord.co/setting-up-a-snowflake-share-with-chord. If you are not using Snowflake, we can still ingest data from your warehouse (Athena, BigQuery, Redshift, Postgres, Databricks, and others) using Parquet with Cloud Storage (S3, etc.). For step-by-step instructions for integrating external cloud storage services to Chord's data warehouse (Snowflake), see Receiving Data from Chord via Cloud Storagehttps://docs.chord.co/sharing-data-with-chord-from-cloud-storage.
Note that Chord's schema expectations and IAM configurations can vary by cloud provider.
The approach
Chord provides two options for sharing data with us.
Option 1: Bring your own models
Send us your data as is. Chord will simply handle the rest. Export your current tables (Orders, Customers, Products, etc.) exactly as they exist in your warehouse today. You do not need to rename columns, change data types, or map relationships.
- Your effort: Low. Simple export.
- How it works: Chord ingests your raw data and will adapt to your schema to generate insights. Chord AI and Copilot understand your schema by reading table and column names, and you can query, analyze, segment, and activate immediately.
- Best for: Teams who want to test the platform quickly without engineering overhead and already have clean, modelled data in a data warehouse.
Note: Some of Chord's data models will not be available if your brand chooses this option.
Option 2: Bring your data and leverage Chord’s models
Map your data to Chord’s Unified Schema, and unlock access to Chord’s full Analytics suite. This option is for teams that want standardized reporting, attribution, and activation across tools.
- Your effort: Moderate. Requires writing SQL transformations.
- How it works: You ensure your data matches our specifications before it leaves your environment. Chord builds the unified tables. You still own the raw data. Chord owns the modelling layer.
- Best for: Teams that want standardized metrics across teams, plug-and-play attribution and lifecycle reporting, and activation-ready datasets.
Required Entities
At a minimum, Chord recommends providing:
- Orders
- Customers
- Line Items
- Products
- Variants
- Payments
- Returns
- Refunds
- Shipments
- Orders
- Order Line Items
- Subscriptions
Setup and automation
When processing a high volume of files from object storage, it’s helpful to name files according to a hierarchy that can be easily queried. Using the following key format is advised:
<source>/<collection>/<year>/<month>/<day>/<hh:mm:ss>/<partition>_<id>.parquet
Here’s a specific example of what this might look like:
example-oms/orders/2025/08/11/20:29:25/2025-08-10-14:00_e791c45e.parquet
This means the file is the orders collection from your custom OMS platform. It was uploaded on 2025-08-11 at 20:29:25 UTC and the data is for the hour of 2025-08-10 14:00. It’s also helpful to include a unique identifier for the file (in this case, shortened GUID e791c45e) corresponding to the compute process that produced this file.