A pipeline target is a number, and a number is not a diagnosis. Pipeline came in at 62 percent of plan. Which part broke? The usual answers are “outbound underperformed” and “marketing sent fewer leads”, both of which are restatements of the same number one level down. The team then adds activity, because activity is the only lever that can be pulled without knowing the answer.
Pipeline engineering is the alternative: treat pipeline generation as a system with declared inputs, a measured yield at every stage and a latency at every handoff, so that a miss names the stage that caused it. This is what to instrument, in what order, and where automation genuinely helps once you can see.
Pipeline as a system, not a target #
Every pipeline motion is the same shape, whatever the channel:
addressable set -> selected -> reached -> engaged -> qualified -> opportunity
Each arrow is a conversion with a yield and a delay. A miss at the end is always attributable to a yield that fell or a delay that grew, somewhere along that chain. Without per-stage numbers, the attribution is a guess, and the guess is usually “more volume”, which is the correct fix for exactly one of the five possible causes.
Three properties make this a system rather than a funnel drawing.
Inputs are declared. The addressable set is a definition, not a list someone exported. It has a rule, the rule is written down, and it can be recomputed. If nobody can say how the target list was produced, the first number in the chain is unfalsifiable and everything after it inherits that.
Every stage has a yield and a latency. Both, always. Yield alone hides the failure mode where conversion is fine and everything takes eleven days, which is the failure that loses deals to whoever answered first. Revenue latency is that cost in full.
Every record carries where it came from. Not channel attribution for reporting. Which rule selected it, which run enriched it, which score qualified it. That is what lets you ask “what was different about the accounts that converted” and get an answer rather than a theory.
Where pipeline actually leaks #
Measured across the chain, the leaks cluster in five places, and only one of them is fixed by more activity.
The addressable set is wrong. The most expensive failure and the least visible, because everything downstream looks like execution. Every yield is depressed by the same factor and each one looks individually explainable. The signal is that the conversions are all mediocre rather than one being bad.
Selection is stale. The set was correct when it was built and is now being worked six weeks later, after the trigger that made those accounts interesting expired. Selection has a shelf life, and most teams never assign one.
The handoff has a queue in it. A lead qualifies on Tuesday and is picked up on Friday. Nothing in the reporting shows this, because both events are recorded and the gap between them is nobody’s metric.
The data was incomplete at the moment of the decision. The account was routed, scored or skipped on a record that was half enriched. It converts poorly and gets counted as a bad-fit account, so the selection rule is tightened and the actual cause survives.
Capacity is silently exceeded. Records are assigned to reps who already hold more than they can work. They are not lost, they are just not touched, and they age out. This appears in the numbers as low conversion rather than as an unworked queue.
Four of those five are invisible to a dashboard that counts stage volumes. They become visible the moment each stage is timestamped and each decision records its inputs.
What to instrument, in order #
Instrument from the end, because the last stage is the one whose numbers people already trust.
| Stage | Yield to measure | Latency to measure | The question it answers |
|---|---|---|---|
| Selected of addressable | Share of the defined set worked this period | Age of the selection rule’s last run | Are we working a current list or an old one |
| Reached of selected | Deliverability, connect rate | Days from selection to first touch | Does the trigger still hold when we arrive |
| Engaged of reached | Reply, meeting booked | Hours from inbound signal to first response | Are we losing to whoever answered first |
| Qualified of engaged | Accepted by sales | Days from engagement to disposition | Is the handoff a queue |
| Opportunity of qualified | Created, and survived 14 days | Days to first meaningful activity | Is capacity the constraint |
Two disciplines make these numbers worth having.
Count records, not events. Sixty emails to twelve people is not sixty touches, and a stage counted in events flatters itself at exactly the moment volume replaces quality.
Snapshot, do not recompute. A stage rate calculated today from a CRM whose fields have since been overwritten is not a historical number. If yield is going to be tracked over time, store the value the stage had when it happened.
Making the chain one system rather than five #
The reason these numbers are hard to get in most stacks is structural: each stage happens in a different tool, so each seam is a place where the trail ends. Enrichment happened in one product, scoring in a warehouse model, routing in the CRM, the touch in a sequencer. Every handoff loses the context of the one before it, which is why “why did this account convert badly” is unanswerable rather than merely unanswered.
When the chain runs as one engine, each stage is a step in a run against the same model, and the run is the instrument:
import { alertConnectorAction, defineAlert, definePlay } from "@cargo-ai/cdk";
export const sourceAndActivate = definePlay("source-and-activate", {
model: accounts,
filter: addressable, // the declared addressable set, in one place
changeKinds: ["added", "updated"],
workflow: pipelineWorkflow, // enrich, resolve, score, route, activate
healthThreshold: 95,
isEnabled: true,
});
export const activationSlowing = defineAlert("activation-slowing", {
schedule: { type: "cron", cron: "0 * * * *" },
scope: { kind: "runs", workflow: sourceAndActivate },
threshold: {
metric: "duration",
aggregation: "p95",
operator: "gte",
value: 900,
},
actions: [
alertConnectorAction({
ref: slack.actions.sendMessage,
config: { channel: "#pipeline", text: "Activation p95 above 15 minutes" },
}),
],
});
The useful part is not that it is fewer tools. It is that healthThreshold makes a degraded batch a condition rather than a discovery, and a latency threshold makes a slowing handoff something the system reports rather than something a quarterly review finds. The stage boundaries sit inside one trace, so a per-record answer to “what did we know when we decided this” exists without anyone reconstructing it.
For the decision layer this sits in, see revenue orchestration. For the practice of keeping selection current rather than batch-built, turning a TAM into prioritized pipeline, continuously.
Where agents fit, and where they do not #
Agents are usually introduced at the stage that is easiest to automate rather than the one that is constraining, which is why the pipeline number often does not move.
Good fits, because the work is judgement over messy inputs and the output is reviewable: researching an account, summarising why it matched, drafting a first touch, classifying a reply, flagging that a signal has gone stale.
Bad fits, because the work is deterministic and an agent adds variance for nothing: assigning by territory and capacity, applying a scoring rubric, enforcing suppression, deciding whether an account is in the addressable set. These are rules. A rule that runs the same way every time is worth more than a rule that reasons.
The ordering rule is simple and ignored often: instrument first, then automate the constraining stage. Automating a stage with a yield of 40 percent produces more of what was not working. The measurement is what tells you which one to touch, which is why it is the first piece of work rather than the reporting layer bolted on afterwards.
Frequently asked questions #
Treating pipeline generation as an instrumented system rather than a target. The addressable set is a declared rule that can be recomputed, every stage from selection to opportunity has a measured yield and a measured latency, and every record carries which rule selected it and which run enriched and scored it. The point is that a miss names the stage responsible instead of producing a debate about effort.
Because a stage can convert well and still lose, by being slow. An inbound lead answered in four hours and one answered in ten minutes show the same conversion definition and very different outcomes, and the gap between qualification and first touch is usually nobody’s metric even though both timestamps exist.
Five places: the addressable set is wrong, so every downstream yield is depressed evenly; selection is stale, so the trigger expired before anyone arrived; the handoff has an invisible queue in it; decisions were made on half-enriched records and the resulting poor conversion gets blamed on targeting; and reps are over capacity, so records are assigned but never worked. Only the first is fixed by more volume.
Per-stage yield and per-stage elapsed time side by side, counted in distinct records rather than events, and snapshotted at the time the stage happened rather than recomputed from fields that have since been overwritten. Most dashboards show stage volumes, which is the one view that cannot distinguish a targeting problem from a speed problem.
Parts of it. Research, summarisation, drafting and reply classification are judgement over messy input and suit an agent. Routing, scoring, suppression and set membership are deterministic rules, and an agent doing them adds variance with no gain. The sequencing matters more than the split: instrument the chain first, then automate the stage that is actually constraining, or you scale the stage that was already working.