Web capture
Each Monday the company's own website, its competitors' pages and recent news about both land in context/ as one pull request: the first run seeds positioning, offerings, an inferred ICP with a disqualifier, competitors, clients and proof; every run after adds a dated note of what changed (pricing, a launch, funding, a customer) and never edits an existing file.
Set up the web-capture cookbook in this project.
1. If the Cargo CLI is not installed yet, follow every step in https://api.getcargo.io/INSTALL.md
2. From the Cargo project, run: cargo-ai cdk add cookbook/web-capture
3. Then follow .claude/skills/web-capture/SKILL.md. Stop for my approval before any paid call, and stop at "cargo-ai project plan" before deploying anything.$cargo-ai cdk add cookbook/web-captureSay this to your agent
Every Monday, read our website, our competitors' pricing and changelog pages, and the news about all of us, and keep context/ current.
Illustrative output #
Fictional records, to show the shape of what comes back.
insight/2026-10-05-web.md
- Northwind: new "Enterprise" tier on /pricing, with SSO and audit logs [R: https://northwind.example/pricing]
- Northwind: raised a $20M Series B led by Contoso Ventures, 2026-10-01 [R: https://news.example/northwind-series-b]
- Northwind: Fabrikam named as a customer, 40% faster onboarding [R: https://northwind.example/customers/fabrikam]
- Tailspin (competitor): Team plan price up from $49 to $69 a seat [R: https://tailspin.example/pricing]
client/fabrikam.md, proof/fabrikam-onboarding.md (reference_permission: unknown)
PR body: proposes adding "Enterprise" to global/offerings.md and the new price to alternative/tailspin.md
A week after the first run: one pull request, four findings across two companies, two new files, two proposed edits, and two reads billed (the pages and the news).
Done when #
DOMAIN,PAGESandCOMPETITORSname the company, its real sections and the competitors worth watching, and the contract passescargo-ai cdk checkprints the agent bound to the repository root, andcargo-ai cdk planreports one agent, two bound connectors, one folder and no model- the first run opened one pull request with the page, competitor and news files, seeded every empty domain, and edited no existing file
- every seeded file carries tags,
icp/names a disqualifier, and everyclient/file carriesreference_permission: unknown - after the baseline was merged, a run with a change opened a pull request adding
insight/<date>-web.mdwith one tagged line per finding, each naming its company, and proposed rather than made any edit to an existing file - a run with nothing worth writing opened no pull request
- after merging and
cargo-ai cdk deploy, a seeded file is readable from the workspace context repository and an agent with thecontextcapability quotes it back with its tag
What it costs #
Fetch live prices before every estimate: cargo-ai connection integration get parallel, plus the
LLM connector’s model. Quote the lookup time with the estimate.
Each run bills two reads: one parallel.extract over every listed URL, the company’s and the
competitors’ (billed per URL, a 404 included), and one parallel.createTask on the lite
processor, the cheapest rung of its price ladder. Plus the harness run, which scales with how much
changed.
Keep the knowledge layer current from the one source every company has on day one: the web, about itself and its competitors. One Claude Code harness agent and its prompt, one pull request a week.
What it does #
- Reads the web, the same way every week. Two
cargo-aicommands generated in the prompt: oneparallel.extractwithfullContentover the company’s own pages and each competitor’s watched pages, and oneparallel.createTasknews question about all of them for the days since the last merged read. A fixednode -eline writes each page tocadence/log/raw/web/pages/<page>.mdorcadence/log/raw/web/competitors/<competitor>/<page>.md, and the news tocadence/log/raw/web/news/<date>.json. No text passes through the agent on its way to disk. - Lets git say what changed.
git diffagainst the default branch shows which pages changed; a news URL not in an earlier file is new. - Seeds once. The first run fills every empty
context/domain:global/,icp/with a disqualifier,alternative/,client/,proof/,signal/. Every sentence tagged[R],[I]or[TR]. - Adds what changed. Every run after writes one dated
insight/<date>-web.md, each line naming its company, plus a file for a newly named customer or competitor. It never edits a file that exists; a change to one is a proposal in the pull request body. - Stops. One pull request, never merged, or none in a quiet week. Nothing here reads a CRM, a call or Slack.
Adds 4 resources and no script.
| File | Resource | Role |
|---|---|---|
infra/agents/web-scribe.ts | defineAgent (claudeCode) | the scribe: model, folder, weekly cron, no env |
infra/agents/web-scribe.prompt.ts | (not a resource) | the whole procedure: domain, pages, competitors, the exact commands, the rules |
infra/connectors/anthropic.ts | defineConnector (anthropic) | the model the harness runs on, billed and metered |
infra/connectors/git.ts | defineConnector (github) | the clone, branch, push and PR path, resolved by binding |
infra/folders/index.ts | defineFolder | the workspace folder the agent is filed in |
Why no script #
The other capture cookbooks read their source with a committed collector, because a fetch loop an
agent re-derives each week silently changes shape. Here the reads are two fixed commands, so the
prompt carries them exactly, and git already does the one thing a collector would add: comparing
this week to the last. What keeps the comparison honest is mechanical: fullContent returns the
whole page rather than excerpts chosen for a question, and a node -e line writes the files from
the command output, so the agent never retypes a page. Both were checked live: two reads of the same
pages a minute apart were byte-identical.
Why competitors are in the default #
A company already knows its own launches, pricing and funding; it made them. What a weekly read
catches that the team does not already know is a competitor changing its pricing or shipping
something. Same two reads, one more set of URLs, and the findings land in insight/ with the
competitor’s name and in proposals against its alternative/ file.
Why the baseline is the last merged read #
The harness clones the default branch, so the only read a run can diff against is one that was
merged. A quiet week opens no pull request and commits nothing, which is why the news window starts
at the date of the last commit under cadence/log/raw/web/ rather than a fixed seven days: the next
run covers the gap. It is also why the first run always opens a pull request, even when every domain
was already seeded.
Why one source #
A cookbook is a prebuilt approach an agent follows. Give it a page and a deal in one run and the
reader cannot tell the two confidence levels apart afterwards. So this cookbook holds one: what is
public, tagged as such. The CRM has win-loss-review, which verifies the ICP this seeds; calls
have call-capture.
Placeholders (edit before deploy) #
DOMAIN,PAGESandCOMPETITORSininfra/agents/web-scribe.prompt.ts. The agent refuses the placeholder domain;COMPETITORSships empty.languageModelon the agent.
What it does not do #
It does not read a CRM, a call, an inbox or Slack; write the workspace context directly; edit a file that exists; write a page by hand; contact anyone; or merge its own pull request.
Verify #
From this skill’s folder:
node --import tsx evals/contract.mjs
From the project root:
npm run check && cargo-ai cdk plan