Your GTM belongs in gitRegister
Blog

Data Architecture for GTM Engineers: What to Learn First

30 Sept
7min read
AurelienAurelien

If you build GTM systems and you are not an engineer, learn data architecture before you learn to code. Your agents already write the code. What they cannot do for you is decide what a record is, how you know two records are the same company, and which system is allowed to change which field. Get those right and a system survives a year of additions. Get them wrong and it becomes the thing nobody on the team wants to touch.

During the fireside at GTM Belongs in Git, someone asked what three skills a non-engineer should go and learn. Noah Adelstein, who runs growth at Netic, gave data architecture as the first. I gave system design. John Kutay, who runs GTM engineering at Rippling, gave identity. They are the same answer from three directions, and this article is that answer made concrete.

Why this, and why now #

A year ago the bottleneck for a GTM builder was writing the pipeline. Now a coding agent writes it in an afternoon, and the bottleneck has moved to what the pipeline writes into.

The failure is predictable. The first version works and everyone is excited. Each addition is built on top of the last one’s assumptions. By the tenth, a company appears three times, an enrichment overwrites a field a rep corrected, and two plays disagree about who owns an account. Nothing is broken in isolation. The shape was wrong from the start, and an agent building faster only gets you there faster.

1. Durable tables: decide what a record is #

Start with the nouns. For most B2B teams the core set is small: companies, people, the relationship between them, signals, and outcomes. Each gets its own table, and each row means exactly one thing.

The test for a durable table is whether it still makes sense in a year. A table called q3_outbound_targets does not. A companies table with a segment column does. Campaign-shaped tables are how a warehouse fills up with forty copies of the same company, each one slightly out of date.

Keep events separate from state. “This person changed jobs on 2026-09-12” is an event, and it goes in a signals table with a timestamp. “This person’s current employer” is state, and it lives on the person. When the two share a row, the history is overwritten every time the state changes, and you lose exactly the record that would explain a score.

2. Keys: one identifier per real thing #

Every table needs a key that identifies one real-world thing and only one. This is the concept that sounds most academic and bites hardest.

For companies, the usual key is the registrable domain, normalized: lowercase, no www., no path. For people, the most stable key in B2B is usually the LinkedIn profile URL, normalized the same way every time. Email is tempting and fragile: people have a work address and a personal one, and the work address dies when they leave.

The rule that matters more than which key you pick: normalize at the door. Every source that writes into a table runs the same normalization before it writes. If one enrichment writes www.acme.com/ and another writes acme.com, you have two companies, and nothing downstream will ever notice.

3. Identity resolution: the same company is not always the same string #

Keys get you most of the way. Identity resolution is the rest, and John called it out as the problem that only reveals itself when you try to run outbound at scale.

The hard cases are predictable, so plan for them:

  • Renames and rebrands. The company changes its name and its domain. Keep a table of previous domains and names pointing to the current record, instead of creating a new one.
  • Subsidiaries. A parent and its regional entities share a brand and have different domains. Decide once whether your model scores the parent, the child, or both, and write the relationship down.
  • People across jobs. A champion who moves companies is one person with two employment records, not two people. That is the signal, and a model that splits them cannot see it.
  • Several identifiers for one person. A LinkedIn URL, a work email and a personal email all resolve to the same person record.

Duplicates are not a cleanup task. They are the output of a missing rule. When you find one, find the rule that would have prevented it and put it at the door.

4. One writer per fact #

System design, for a GTM builder, mostly comes down to a single question asked about every field: who is allowed to write this?

When enrichment, a rep and a sync all write job_title, the value is whichever one ran last. Decide an owner per field. Enrichment owns firmographics. Reps own qualification notes. The scoring model owns the tier and nothing else writes it. Other systems read.

This is also what makes agents safe to run. An agent that knows it may write the tier and nothing else can be given the job without anyone watching every run. One allowed to write anywhere needs a human reading everything it does, which defeats the point.

What this looks like in a week #

You do not need a data engineering degree. You need a short document and the discipline to keep it true:

  1. List your core tables and one sentence on what a row means.
  2. Write the key for each, and the normalization rule for it.
  3. List the identity cases you have already hit, and the rule for each.
  4. For every field that more than one system touches, name its owner.

Put it in the same repository as the pipelines, so an agent reads it before it builds the next one. That single file prevents more rework than any amount of prompt tuning.

Engineer or GTM person? #

The question came up at the event, and the answers held up. John’s was that it depends on stage and on how many internal teams consume the data: one product and one sales team does not need a data engineering function. His warning ran the other way too. An experienced data engineer who is not excited by go-to-market problems will tend to optimize the pipeline instead of the pipeline’s purpose. My answer, which John agreed with: if you have the budget, hire one person from each side and make them work together. If you do not, pick the curious GTM person and hand them this list.

For the wider picture of the role, see the technical GTM engineer. For how the objects fit together in a CRM, see CRM data architecture.

FAQ #

AurelienAurelienSept 30, 2026

Give your agents a runtime

Bring the agents you have.Start free, deploy in one command.