Most GTM agents fail on context, not on reasoning. The model is fine. It can write, summarise, classify and follow a plan. What it cannot do is know that your ICP excluded agencies in March, that the competitor the prospect just named loses on data residency and wins on price, or that this account already heard from a rep eleven days ago. Without those, it produces copy that is fluent, plausible and indistinguishable from what any other vendor would send.
“Give it more context” is not an instruction anyone can act on, because context is four different things with four different failure modes. This is what each one is, how to give an agent each one as code, and what keeps them true after deploy.
The four things “context” means #
| Layer | The question it answers | Where it lives | How it goes wrong |
|---|---|---|---|
| Knowledge | What is true about our company and market? | Prose the company wrote: ICP, positioning, objections, proof | Stale, or asserted with no evidence behind it |
| State | What is true about this account, right now? | The data model: accounts, contacts, usage, deals | Fresh in one system, hours old in the agent’s copy |
| Tools | What is the agent allowed to do about it? | A typed action surface | Untyped, ungated, or so wide the agent picks wrong |
| Memory | What happened last time? | Run history, allocation history, prior sends | Absent, so the agent repeats itself to the same person |
A prompt can carry a little of layer one. It cannot carry any of the other three, which is why prompt engineering plateaus and why the interesting work is infrastructure. For the layer underneath all of this, see what GTM infrastructure is.
Knowledge: write it down, version it, make it citable #
Company knowledge is the layer teams skip, because it is the only one that is not a technical problem. Nobody has written the ICP down in a form an agent can read. It exists in a deck, in three people’s heads and in a Notion page last edited in January.
The fix is unglamorous. Put it in markdown, put the markdown in git, and deploy it as a resource the agent reads:
import { defineContext } from "@cargo-ai/cdk";
export const context = defineContext({
path: "./context",
});
Every .md and .mdx file under that directory is synced into the workspace’s context repository on deploy. One repo per workspace, committed, so a change to the ICP is a diff a person reviews rather than an edit nobody sees.
Three properties matter more than the file format.
Claims carry a confidence level. A sentence like “prospects churn because onboarding takes six weeks” is either something one customer said once or something six said independently, and those are different facts. Marking each claim hypothesis, validated or proven, and permitting outbound copy to cite only the last two, is the cheapest guard against an agent stating an anecdote as a finding. It costs one frontmatter field.
Knowledge is separated from instruction. The system prompt says how to behave. The context repo says what is true. Merging them means every factual correction is a prompt change, and every prompt change risks the behaviour. Keep the prompt short and let it point at the knowledge.
Files cross-reference each other. An objection file that names the competitor it belongs to, a persona that names its ICP, a proof point that names the claim it supports. Retrieval over a flat pile of documents returns the document that lexically matches. Retrieval over a graph returns the document and the two it depends on.
State: give the agent the model, not a screenshot #
Pasting account data into a prompt is a snapshot. By the time the agent acts, the deal moved, the usage spiked or the account was already assigned.
The alternative is to let the agent read the same unified model the rest of the engine reads, as one entry in what it is allowed to use:
import { defineAgent } from "@cargo-ai/cdk";
export const accountResearcher = defineAgent("account-researcher", {
connector: openai,
languageModel: "gpt-4o",
systemPrompt: [
"You research accounts for the EMEA enterprise team.",
"Cite the context repo for any claim about our product or market.",
"Never assert a fact you did not read from a tool or a model.",
].join("\n"),
uses: [
{ ref: accounts, readOnly: true },
{ ref: contacts, readOnly: true },
enrichAccount,
hubspot.actions.searchRecords,
],
});
Two details in that block are doing most of the work. readOnly: true on the models means the agent can ask and cannot write, which is the correct default for anything that reads customer records. And uses is one array covering models, tools, connector actions and sub-agents, so the whole surface the agent can reach is readable in one place rather than spread across a UI.
The unified model is what makes this worth doing. An agent reading five connector APIs has to resolve identity itself, and it will do it badly: the same company under two domains becomes two accounts, and the agent writes to both. An agent reading one model where that resolution already happened inherits it.
Tools: type the actions, gate the irreversible ones #
An agent that can call anything will eventually call the wrong thing. The useful constraint is not a longer prompt, it is a narrower and better typed surface.
Three rules hold up in practice.
A tool is a workflow, not an API call. “Enrich this account” is a tool. “POST to this endpoint” is not, because the agent then owns the retry policy, the provider fallback and the error handling, and it owns them badly. Wrapping the whole chain as one typed tool means the agent chooses intent and the infrastructure chooses mechanics.
Reads are wide, writes are narrow. Let the agent read everything it might need. Let it write to a small, named set of places, and make each write reviewable after the fact.
Anything that reaches a person is gated. Drafting an email is an agent decision. Sending it is a policy decision. Keeping those separate is what makes the rest of the autonomy safe to grant, and it is the same reason a deploy has a review step.
For what a platform needs in order to support this at all, see the agent-ready GTM platform. For when to reach for an agent versus a deterministic workflow, multi-agent workflows for B2B revenue teams.
Memory: the run history is the memory #
Agent memory is usually discussed as a vector store. For GTM the more valuable memory is boring and already exists: what this engine did, to whom, and when.
The questions that actually matter are historical. Has anyone at this account been contacted in the last thirty days? Which rep holds it, and why them? Did we already try this angle and get nothing? An agent that cannot answer those will re-contact a live opportunity, and a single instance of that costs more trust than the agent saves in a quarter.
So the memory to wire in first is the engine’s own trace: run history keyed to the record, allocation history for assignments, and a ledger of what was sent. Retrieval over past conversations is useful on top of that. It is not a substitute for it.
Keeping context true after deploy #
Context rots in a specific direction. It is written once, when someone cares, and then the product ships a feature, a competitor changes its pricing, and the ICP quietly narrows. Nothing fails. The agent keeps producing fluent output built on a market that stopped existing.
Four habits keep it honest, and none of them is a tool purchase.
Date everything absolutely. “Recently” and “next quarter” are unreadable six months later. A claim with a date can be judged stale. A claim without one cannot.
Make new knowledge a pull request. Something learned on a call becomes a diff against the context repo, reviewed by a person, merged. That is slower than editing a doc and it is the only version that leaves a record of who agreed to it.
Archive rather than overwrite. When positioning changes, keep the previous version. The question “what were we claiming in March” is asked more often than anyone expects, usually when a deal from March comes back.
Evaluate output against the context, not against taste. The check that catches drift is not “does this read well”, it is “is every factual claim in this draft traceable to a file”. An agent that cites is an agent whose mistakes are findable.
Frequently asked questions #
Four separate things. Knowledge, meaning prose about your ICP, positioning, objections and proof. State, meaning live data about the account in front of it. Tools, meaning the typed set of actions it is allowed to take. And memory, meaning what the engine already did to this account and when. A prompt can carry a little of the first and none of the other three, which is why context is an infrastructure problem rather than a writing problem.
In version-controlled markdown that deploys into the workspace as a resource, not in the system prompt and not in a document nobody reviews. Keeping it separate from the prompt means a factual correction does not risk the agent’s behaviour, and keeping it in git means every change to what the company claims is a diff a person approved.
Constrain the claims it may make to ones written down, and make each written claim carry its evidence and its confidence. A useful rule is that outbound copy may cite only claims heard independently at least twice, with anything heard once marked as a hypothesis and excluded. Then check drafts for traceability rather than for tone: every factual sentence should point at a file.
A unified model. An agent reading several source systems has to resolve identity itself, and it will treat one company under two domains as two accounts, then write to both. A unified model has already done the resolution and the deduplication, so the agent inherits one answer instead of inventing one.
Wide on reads, narrow on writes, and gated on anything that reaches a person outside the company. Drafting, researching, scoring and summarising can run unattended because they are reversible. Sending, spending and changing a customer’s record are not, and a human approving those is what makes the unattended half safe to grant.
Retrieval finds documents that resemble the question. Memory answers what this system already did: who was contacted, when, by which rep, with what result. For go-to-market the second is the one that prevents visible mistakes, because the common failure is not a vague answer, it is emailing someone your colleague emailed last week.