A GTM stack does not fail because data is missing. It fails because the same fact disagrees with itself in four systems. The warehouse says the account has 340 employees. Salesforce says 250 because someone typed it in 2024. The enrichment tool says 410. The sequencer personalises on whichever one it received last. Nobody is wrong and the output is still wrong.
Revenue data middleware is the layer that decides which of those is true, keeps it consistent as it moves, and leaves a record of every decision. This is what it has to do, how it differs from the syncs most teams already run, and when the honest answer is that you do not need it yet.
What sits in the middle #
Most revenue stacks have four tiers and nothing between the last two:
- Sources. Product events, billing, the website, enrichment providers, intent feeds.
- A store. A warehouse, or a set of connector tables, or both.
- Systems of action. CRM, sequencer, ads, support, Slack.
- People and agents, working in tier three.
Data gets from one to two through ETL, which is a solved and well-tooled problem. Getting from two to three is where teams improvise: a reverse ETL job for one field set, a Zap for another, a nightly script somebody wrote, and a CRM workflow filling the gaps. Each is defensible. Together they are an undocumented integration layer with no owner.
Middleware is the decision to make that layer explicit instead of emergent.
Why point-to-point stops working #
The failure is combinatorial and it arrives suddenly. Four systems that each need to know about accounts is six pairs. Seven systems is twenty-one. The count of things that can silently disagree grows faster than the team maintaining them.
Three specific costs show up before anyone names the pattern.
Every integration re-solves identity. Is acme.com the same account as acme.io? Is the contact who signed up with a personal address the same person as the one on the opportunity? Each pipeline answers separately, and each answer is slightly different, so the account count differs in every system.
Schema changes break silently downstream. A column renamed in the warehouse fails one job loudly and leaves three others writing nulls. Nothing alerts, because writing a null is a successful write.
Nobody can answer why a record looks like this. The field says “Enterprise”. Which job set it, from which source, at what time, under which rule? With point-to-point syncing, that question requires reading four codebases and two UIs, so in practice it is not asked and the field is not trusted.
The six jobs of a middleware layer #
A sync moves rows. Middleware does six things a sync does not, and the gap between those lists is the whole argument for having one.
Identity resolution, once. One place decides that these rows are one account and those contacts belong to it. Everything downstream inherits that decision rather than repeating it. This is the single highest-value property, because it is the one that changes the numbers rather than the plumbing.
A schema contract. Downstream systems depend on a declared shape, not on whatever the source happens to emit this week. A change to the contract is a change someone reviews, and a source that stops honouring it fails at the boundary instead of three systems later.
Conflict rules. When the warehouse and the CRM disagree, something must decide. Precedence by source, by recency, or by field, declared explicitly. Most stacks have this rule and it is implicit in the order the jobs happen to run, which means it changes whenever a schedule does.
Idempotency. Writes are keyed, so re-running a job does not create a second account or send a second email. Without it, every retry is a risk and so nothing gets retried, which is worse.
Back-pressure and rate limits. Every destination API has limits. Handling them in one place means a spike in the source is absorbed rather than turned into a partial write across six systems.
A trace per record. For any field on any record: which run wrote it, from which input, under which version of the logic. This is what turns “the data is wrong” from an argument into a lookup.
Middleware, reverse ETL and iPaaS #
These overlap enough that vendors describe themselves with all three words. The distinctions that matter to a buyer:
| What it optimises for | Where it stops | |
|---|---|---|
| Reverse ETL | Moving warehouse tables into SaaS tools reliably | The warehouse has to already hold the right answer, and the logic lives in dbt models upstream |
| iPaaS / workflow tools | Connecting anything to anything quickly | No shared data model, so identity and conflicts are re-solved per flow |
| Revenue data middleware | One resolved model plus governed writes into systems of action | Not a replacement for ETL into the warehouse, and not a CRM |
Reverse ETL is the right tool when the warehouse is genuinely the source of truth and the job is delivery. It becomes the wrong shape when the logic that produces the answer needs data the warehouse does not hold, such as a live enrichment call or a rep’s current capacity, because then the decision is being made in two places.
Declaring the layer instead of assembling it #
The practical version of “make it explicit” is that the model, the relationships between models and the destinations are declared in one project, type-checked, diffed before deploy and reviewed in a pull request:
import { defineModel, defineRelationship, defineSegment } from "@cargo-ai/cdk";
export const accounts = defineModel("accounts", {
connector: hubspot,
extractSlug: "fetchRecords",
});
export const contacts = defineModel("contacts", {
connector: hubspot,
extractSlug: "fetchRecords",
});
export const contactsToAccounts = defineRelationship("contacts-to-accounts", {
from: { model: contacts, column: "company_domain" },
to: { model: accounts, column: "domain" },
relation: "manyToOne",
});
export const expansionCandidates = defineSegment("expansion-candidates", {
model: accounts,
filter: {
conjonction: "and",
groups: [
{
conjonction: "and",
conditions: [
{
kind: "string",
columnSlug: "lifecycle",
operator: "is",
values: ["customer"],
},
{
kind: "number",
columnSlug: "seats_used",
operator: "greaterThan",
value: 40,
},
],
},
],
},
});
What that buys is not elegance. It is that the join between contacts and accounts is written down once, in a file, instead of being re-expressed in every job that needs it. When it changes, one diff changes it everywhere, and the diff is the review. This is the same argument as GTM as code applied to the data layer specifically, and it is why unifying customer data is a prerequisite rather than a later phase.
How to tell whether you need this #
You do not need a middleware layer to run two syncs. The threshold is usually reached at a recognisable point, and these are the signals in the order they tend to appear.
- Two systems report different counts for the same segment, and reconciling them is a recurring meeting.
- A field nobody trusts has a manual override process built around it.
- Onboarding a new tool takes weeks because its data has to be wired from four places.
- An outage or a schema change caused wrong records to be written, and nobody could say how many.
- Someone asks “why did this account get routed to that rep” and the answer takes a day.
One of those is normal. Three of them means the integration layer already exists, undocumented, and the choice is whether to name it.
The cheapest first step is not a purchase. Pick the one entity that causes the most arguments, usually the account, decide in one place how it is identified and which source wins per field, and write that down. Most of the value in this layer comes from that single decision being made explicitly rather than from any tool that enforces it.
Frequently asked questions #
The layer between a warehouse or data store and the systems revenue teams act in: CRM, sequencer, ads, support. It resolves identity once, holds a declared schema contract, decides which source wins when two disagree, writes idempotently, absorbs rate limits, and records which run set which field. A sync moves rows; middleware decides what is true before the rows move.
Reverse ETL delivers warehouse tables into SaaS tools and assumes the warehouse already holds the right answer. Middleware is where the answer is produced, which matters when the decision needs something the warehouse does not hold, such as a live enrichment result or a rep’s current load. Teams often run both: ETL in, middleware to decide, reverse ETL or direct writes out.
An iPaaS connects systems quickly and gives you no shared data model, so every flow re-solves identity and conflict resolution on its own. That is fine for a handful of flows. It stops being fine at the point where two systems disagree about the same account and no single place decides which is right.
Reverse ETL products such as Census and Hightouch do delivery in that direction. iPaaS tools such as Workato, Tray and n8n do it as general integration. GTM infrastructure platforms such as Cargo do it as part of a unified model where the sync, the logic that decides the value and the action that follows are steps in one run. The question to ask a vendor is not whether it supports both systems, it is whether a single trace shows why a specific field on a specific record holds the value it holds.
Identity. Two systems end up with different account counts because each pipeline decided separately whether two domains are one company. Everything downstream inherits that split: routing sends one buying committee to several reps, reporting double counts, and personalisation cites the wrong company size.