Your scoring model is wrong somewhere right now, and nobody is looking for where. The fix is a scheduled job that looks for you: it reads pipeline data, finds accounts the model is suppressing, works out which rule did it, and opens a pull request that changes the rule and shows how many accounts it moves. A human reads the preview and merges, or does not.
Cole D’Ambra runs this at Plain and walked through a live one on stage in San Francisco on 2026-09-24. This is how it works, and what you need in place to build your own.
The PR that started it #
A large security company was all over Plain’s website: an intent score of 217, 17 sessions, over an hour on the site. It was also marked out of the addressable market, so its card read cold and nobody was alerted.
Both facts were true. Plain’s intent platform only raises an account to act-now if it is inside the ICP, and this one was not. The reason was a rule written in 2025: at least 30% of a company’s staff had to be engineers. This company was a very large enterprise at 24.5%.
The weekly job caught it, checked whether the rule still made sense, found the same rule was muting other accounts too, and opened a pull request. The diff was three lines: the engineering density threshold went from 30% to 20%. Under it sat the preview, computed against all 55,000 accounts in the CRM:
| Move | Accounts |
|---|---|
| Out to tier 3 | ~200 |
| Tier 3 to tier 2 | 86 |
| Tier 2 to tier 1 | 23 |
It also named the accounts moving in. Cole’s description of the list: the who’s who of the billboards near his apartment. He merged it. A few hundred accounts re-scored overnight, the security company’s card went hot, and the AE and SDR got the alert.
The loop, step by step #
1. Read one place. The job reads from a single store of GTM data. Plain calls theirs the Plain brain: about 15 tables holding call transcripts, channel attribution, intent signals and search performance. If the evidence lives in six tools, the job has six integrations to break.
2. Ask two narrow questions. Not “how can scoring improve?” Two questions with yes-or-no answers:
- Are we missing high-value accounts, either in new qualified pipeline or in high-intent accounts the model marked out of market?
- Did we miss a signal, something a prospect said publicly or on a call that we should have found first?
3. Judge the rule, then its reach. A yes triggers an LLM judgment call: given this account, does its classification still make sense? If not, the job checks whether the same rule is wrong for other accounts. One misfire is an exception to note. A pattern is a rule to change.
4. Propose the smallest diff. The change lands as a pull request against the file that holds the rule, usually one named threshold. A small diff is what makes the next step cheap to review.
5. Price it before anyone approves. The job runs the modified model against every account and reports tier transitions and example accounts. This is the part a UI-configured model cannot give you, and it is what turns “should we loosen this?” from an opinion into a table.
6. A human says yes. Always. The agent finds and prices. The owner of the model decides.
What has to be true before you build it #
The loop is short. The prerequisites are where teams stall.
The model is code. Plain’s scoring model is one TypeScript file with over 500 automated checks behind it in GitHub Actions. Moving it there took under a day. If your model lives in CRM formula fields, start there, because an agent cannot open a pull request against a settings screen.
Thresholds have names. ENGINEERING_DENSITY_MIN = 0.3 can be found, changed, reviewed and blamed. A 0.3 typed into a filter cannot.
Scoring is replayable. The preview needs to run the model on every account without side effects: same inputs, same tier, nothing written. If scoring and alerting are one step, split them first.
// Pure: accounts in, tiers out. No writes, so a preview can run it at will.
export const ENGINEERING_DENSITY_MIN = 0.2;
export function tier(account: Account): Tier {
if (account.engineeringDensity < ENGINEERING_DENSITY_MIN) return "out";
// ...the rest of the model
}
The preview is a script, not a promise. Check out the base branch, score everything, check out the PR branch, score again, diff the two. Post the counts and a sample of named accounts as a PR comment. It is a short script, and it is the most persuasive review artifact you will ever put in front of a sales leader.
Downstream reads the tier, not the rule. After Plain’s merge, alerts, the account brief and buying-committee sourcing all followed from the new tier on their own: surface the likely committee, rank it, find and validate contacts, and hand reps a Slack alert with the brief attached. If every consumer re-implements the rule, one merge fixes one of them.
Keep it honest #
- Never auto-merge a scoring change. Every play downstream inherits it. The preview makes review fast, not optional.
- Cap the rate. Weekly is enough. A job that opens a PR every day trains the owner to stop reading them.
- Log the misses you do not fix. A single out-of-ICP account with high intent is sometimes just noise. Writing down “seen, left alone, because” is what lets next month’s job tell a pattern from a one-off.
- Watch capacity. Moving 300 accounts up a tier is 300 more alerts. Cole held Plain’s intent alerts to three per rep per day until the platform was running durably, and only then raised it to 10 or 15.
Where this leaves the owner #
Before, a change to the model was a rebuild: one more signal, the whole thing rewritten, and nothing to compare it to. Cole rebuilt five times before he saw the pattern. After, the model improves in small reviewed steps, most of them proposed by the job, each with its consequence attached. The owner’s job moves from writing every rule to deciding which proposed ones are right.
That is what GTM as Code buys you once it is in place: not just a history of changes, but a system that proposes the next one.
FAQ #
Three things: the scoring model as code in a repository, a single store of pipeline and intent data it can query, and a way to run the model on every account without side effects so it can preview a change. With those, a scheduled job can find suppressed accounts, identify the rule responsible, and open a pull request.
A report attached to a proposed scoring change that shows how many accounts move between tiers, and which ones, if the change is merged. It is computed by scoring every account under the current model and under the proposed one, then comparing the results.
No. Every downstream play, alert and audience inherits the scoring model, so a wrong change repeats itself on every account it touches. The agent’s job is to find and price the change. A human who owns the model approves it.
Weekly is a good default. It is frequent enough to catch a suppressed account before the buying window closes, and infrequent enough that the owner still reads every pull request carefully.