RevGuild Practices

Score a 4,000-account book with a written rubric you can argue with

Anthropic's Travis Bryant scored 4,000 accounts overnight against two written rubrics, tuned on a test territory first. How it works, and where it breaks.

The problem

Territory planning has a volume problem that care does not fix. A mid-market book of a few thousand accounts is too large to hold in one head and too consequential to rank by feel — so it gets ranked by feel anyway: by recency, by logo size, by whoever a rep spoke to last quarter. That order decides where a year of effort goes, and nobody can reconstruct why any given account sits where it does.

The structural shape underneath: the priority order carries no reason. An order with no reason attached can be accepted or ignored, but it cannot be corrected. That is why both obvious fixes fail against it.

Adding a process step — a quarterly review, a planning offsite — produces a list that was argued into existence in a room and loses the argument the moment the room empties. Three weeks later nobody remembers why account 340 outranked account 12. The reasoning existed; it never became an artefact.

Buying a tool fails the other way. A fitted propensity model — weights learned from historical conversions — scales to the whole book and emits something that looks defensible, but its reasoning is not inspectable by the people who have to act on it. A rep who disagrees with a model score has no move except to ignore it. So they do, and the book gets worked by feel again, now with a dashboard on top.

What the job actually needs is a ranking that scales and can still be argued with.

How it works

Travis Bryant, Head of US Mid-Market GTM at Anthropic, described his annual territory-planning run in a post on the Claude blog published in May 2026. What follows is his account of it, in execution order.

1. Write the rubric first — and write more than one. Bryant calls the project account propensity scoring, and says he started by defining two five-dimension scoring rubrics with Claude: one for tech accounts and one for industries.

2. Name the dimensions explicitly. The tech rubric scores agent opportunity, internal transformation, AI commitment, white space against existing spend, and industry fit. The industries rubric is a different rubric; the post names knowledge-worker density and public AI commitments measured by mentions on the company’s open jobs page, but does not list all five of its dimensions.

3. Tune the weights on a test territory before running the rest. Bryant describes the pattern as: tell Claude what dimensions to score on, run a test territory, check the output, adjust the weights, run the next territory. He notes that none of the prompts were technical, and gives the register of a weight adjustment directly:

I think D4 is probably weighted a little heavy; bring it down a bit

Travis Bryant source ↗

4. Run the full book unattended. Claude Cowork ran overnight, scoring each account one by one across the 4,000-account list, drawing on deep web research, Salesforce data and BigQuery data. He says he did it in one night. His point of comparison is his own history rather than a measurement of this project: in previous companies and roles, he says, work like this ran for hundreds of hours across RevOps, FP&A and marketing.

5. Emit a rationale, not just a number. Every account gets a numerical score and a written rationale for every dimension.

6. Ship it as something a seller opens. The output is an interactive dashboard: accounts ranked by score within each territory, the per-dimension rationale, and — on hover — potential use cases and comparable case studies for prospecting.

And on the interface he uses for Claude Cowork work generally:

I tried Claude Code, but never got comfortable with working with the terminal.

Travis Bryant source ↗

What the post does not specify: how Claude Cowork was connected to Salesforce and BigQuery; whether scores are refreshed between the annual planning runs; whether anyone else on the team runs the same rubrics; and — the important omission — any accuracy figure, win rate or revenue outcome for the resulting ranking. There is no reported check of the scores against closed-won.

Why it works — our read

The practice changes exactly one thing at the input: the unit of work stops being a score and becomes a named dimension with a weight on it. Everything downstream is a consequence of that swap.

Because the unit is a named dimension, an error has an address. When a rank looks wrong, the question stops being do I trust this? and becomes which dimension fired, and is it weighted too heavily? That second question is answerable by a sales manager with no data science, and answering it edits the artefact rather than producing a complaint about it. An opaque score offers no equivalent move: the only available responses are compliance and quiet non-compliance.

Because every dimension emits a written rationale, checking becomes cheaper than redoing — and that ratio is what makes the volume tractable at all. Reading a paragraph and disagreeing with it costs a minute; reconstructing why an account deserves attention costs an hour. While review stays an order of magnitude cheaper than the original work, one person’s judgment can cover four thousand accounts. The number is a sort key. The rationale is the product.

Because the weights were tuned on a test territory first, the blast radius was bounded during the window when the thesis was most likely to be wrong. This is staged rollout — an old engineering idea — applied to judgment rather than to code. A rubric wrong across one territory costs an afternoon. The same rubric wrong across the whole book seeds every downstream conversation with a bad reason, and bad reasons are much harder to un-ring than bad numbers.

And because the tuning happened where the thesis-owner could operate himself, the hand on the dial belonged to the person with the market view. Put a terminal in between and you have inserted a translator into the one loop that has to stay tight.

The transferable principle is one line: encode the judgment, not the answer. The counterfactual is the same book scored by a fitted model — a ranking arrives, nobody can interrogate it, and the interrogation was the work.

Where it breaks

The post does not discuss failure modes; this is our assessment.

  • The rubric encodes one person’s theory of the market, then applies it at full confidence to every account. Intrinsic to the practice — and the rationale makes it worse, not better. A wrong ranking with an articulate explanation is harder to dislodge than one with a bare number. Symptom: every rationale reads well and none surprises you.
  • The tuning loop has no ground truth in it. Nudging a weight down because it feels heavy calibrates against the author’s expectation, not closed-won. Nothing in the published account closes that loop. Symptom: a year later nobody can say whether the weights were right.
  • Proxy dimensions decay. Counting AI mentions on a jobs page discriminates only while posting AI roles is a choice. Symptom: a dimension where nearly every account scores high, consuming weight and adding no ordering.
  • A snapshot ages. Scores from one overnight run reflect that night’s research, and the post reports no refresh between the annual planning runs. Symptom: a rep opening a call with a rationale two quarters stale.
  • Segment rubrics multiply. Two is a morning of maintenance a year. Eleven is somebody’s second job.

Present the output as a prioritisation aid, not a validated model. And don’t do this at all if you cannot name your dimensions before you open the tool: ask a model to propose them and you adopt a theory of your market that nobody chose, and never notice.

What has to be true at your company

Somebody owns the market thesis and will write it down. This practice does not generate a point of view; it scales one. If no named person will put five dimensions on paper and be wrong about them in public, you get a rubric nobody defends.

The book is big enough that ranking is a real problem. Below a few hundred accounts a team holds the whole book in working memory, and the rubric costs more than the ordering is worth.

Account-level data exists, is queryable, and is fresh enough to score on. Bryant’s run used web research plus Salesforce and BigQuery; the shape that matters is one system of record for pipeline and one for usage or spend. Dimensions keyed to a field reps update sporadically will produce confident sentences built on stale data.

Someone owns territory assignment and will act on the output. A ranking that changes neither territory boundaries nor sequencing changes nothing.

The culture tolerates a leader publishing an opinionated ranking that reps will contest — and has a forum where contesting it lands somewhere.

Across company shapes: at eight people, run one rubric over two hundred accounts in a spreadsheet, founder writing the dimensions — the value is forcing the thesis into words. At eight hundred the risk inverts: the rubric stops being a working document and becomes policy, so version it, date it, name its owner. An unowned rubric outlives the thesis it encoded and keeps ranking the book against a market that has moved.

Try it this week

Take one rep’s territory — fifty accounts is plenty. Before opening any tool, write five dimensions on paper with a weight for each, summing to 100. Twenty minutes, one person, no approval needed. Then score ten accounts against it by hand, writing one sentence of rationale per dimension. Another thirty minutes. That is the whole first move: no automation yet.

What you should see if it is working: at least two of the ten move materially up or down against the current gut ranking, and you can name the dimension that moved them. Both halves matter — movement without a nameable cause is noise.

How you would know it is not: the ten come out in roughly the order you would have written from memory. Then the rubric is restating your priors rather than testing them, and automating it will scale a ranking you already had. Fix the dimensions or stop — do not scale it and hope.

Only once ten hand-scored accounts have told you something you did not already believe is it worth automating the rest of the territory.

Sources

  1. How an Anthropic sales leader uses Claude Cowork to run a 4,000-account book Anthropic · article · published May 20, 2026 · accessed Aug 3, 2026 · primary

Travis Bryant — Head of US Mid-Market GTM, Anthropic — role and organisation as stated in the cited source at the time it was published. Quotes are verbatim from the linked source; everything under “Why it works — our read” and “Where it breaks” is RevGuild analysis, not a claim by Travis Bryant. Spotted something wrong? hello@revguild.org — we correct in place and say so.

Frequently asked questions

Isn't this just lead scoring with extra steps?

Same goal, opposite failure mode. Classic lead scoring fits weights to historical conversions and hands you a number whose reasoning you cannot inspect. This starts from dimensions a human wrote down on purpose and emits a written rationale per dimension, so the ranking can be disagreed with specifically. The trade is real: you give up statistical fit and get auditability instead.

Why not just have our data team fit a statistical propensity model?

Build it if you have the conversion volume to fit one and the appetite to maintain it. Note that Bryant calls his own project account propensity scoring — the contrast we are drawing is not with the label but with the mechanism: weights learned from historical conversions versus weights a human chose and wrote down. A rubric is what you reach for when you have a market thesis and not enough closed-won data to validate it, which is the normal condition for a new segment or a new product. They are also not exclusive: a rubric is a reasonable way to state the hypotheses a fitted model would later test.

Our CRM data is a mess. Does this still work?

Partly. Bryant's run drew on web research alongside Salesforce and BigQuery, so dimensions grounded in public signal degrade gracefully when internal data is patchy. Dimensions that depend on internal fields — white space against existing spend, for example — inherit whatever mess is in the CRM, and the rationale will state a confident reason built on a stale field. Score the dimensions you actually have data for.

Won't reps game the rubric once they know the dimensions?

They can only game it if scoring depends on rep-entered fields. Dimensions computed from public research and system-of-record spend are hard to inflate from a CRM form. The bigger risk is not gaming but anchoring: once the ranking is published, people stop noticing accounts that the rubric never had a dimension for.

Do we need Claude Cowork specifically?

Nothing in the mechanism requires that product — dimensions, weights, a test territory and a written rationale are tool-agnostic. What the source does suggest is that the interface matters: Bryant says he never got comfortable working in a terminal. If the only person who can turn the dial is an engineer, the person with the market thesis is no longer the person tuning it.

How often do we have to re-score?

The source says every account in the book needs a score every fiscal year, and describes territory and prospect-list work as a quarterly rhythm. What it does not report is whether scores are refreshed between those planning runs. The practical constraint is that scores age at the speed of their fastest-moving dimension — if one dimension keys on hiring signals, the ranking is stale within a quarter, whatever the planning calendar says.

Apply →