The problem
Territory planning has a volume problem that care does not fix. A mid-market book of a few thousand accounts is too large to hold in one head and too consequential to rank by feel — so it gets ranked by feel anyway: by recency, by logo size, by whoever a rep spoke to last quarter. That order decides where a year of effort goes, and nobody can reconstruct why any given account sits where it does.
The structural shape underneath: the priority order carries no reason. An order with no reason attached can be accepted or ignored, but it cannot be corrected. That is why both obvious fixes fail against it.
Adding a process step — a quarterly review, a planning offsite — produces a list that was argued into existence in a room and loses the argument the moment the room empties. Three weeks later nobody remembers why account 340 outranked account 12. The reasoning existed; it never became an artefact.
Buying a tool fails the other way. A fitted propensity model — weights learned from historical conversions — scales to the whole book and emits something that looks defensible, but its reasoning is not inspectable by the people who have to act on it. A rep who disagrees with a model score has no move except to ignore it. So they do, and the book gets worked by feel again, now with a dashboard on top.
What the job actually needs is a ranking that scales and can still be argued with.
How it works
Travis Bryant, Head of US Mid-Market GTM at Anthropic, described his annual territory-planning run in a post on the Claude blog published in May 2026. What follows is his account of it, in execution order.
1. Write the rubric first — and write more than one. Bryant calls the project account propensity scoring, and says he started by defining two five-dimension scoring rubrics with Claude: one for tech accounts and one for industries.
2. Name the dimensions explicitly. The tech rubric scores agent opportunity, internal transformation, AI commitment, white space against existing spend, and industry fit. The industries rubric is a different rubric; the post names knowledge-worker density and public AI commitments measured by mentions on the company’s open jobs page, but does not list all five of its dimensions.
3. Tune the weights on a test territory before running the rest. Bryant describes the pattern as: tell Claude what dimensions to score on, run a test territory, check the output, adjust the weights, run the next territory. He notes that none of the prompts were technical, and gives the register of a weight adjustment directly:
I think D4 is probably weighted a little heavy; bring it down a bit
4. Run the full book unattended. Claude Cowork ran overnight, scoring each account one by one across the 4,000-account list, drawing on deep web research, Salesforce data and BigQuery data. He says he did it in one night. His point of comparison is his own history rather than a measurement of this project: in previous companies and roles, he says, work like this ran for hundreds of hours across RevOps, FP&A and marketing.
5. Emit a rationale, not just a number. Every account gets a numerical score and a written rationale for every dimension.
6. Ship it as something a seller opens. The output is an interactive dashboard: accounts ranked by score within each territory, the per-dimension rationale, and — on hover — potential use cases and comparable case studies for prospecting.
And on the interface he uses for Claude Cowork work generally:
I tried Claude Code, but never got comfortable with working with the terminal.
What the post does not specify: how Claude Cowork was connected to Salesforce and BigQuery; whether scores are refreshed between the annual planning runs; whether anyone else on the team runs the same rubrics; and — the important omission — any accuracy figure, win rate or revenue outcome for the resulting ranking. There is no reported check of the scores against closed-won.
Why it works — our read
The practice changes exactly one thing at the input: the unit of work stops being a score and becomes a named dimension with a weight on it. Everything downstream is a consequence of that swap.
Because the unit is a named dimension, an error has an address. When a rank looks wrong, the question stops being do I trust this? and becomes which dimension fired, and is it weighted too heavily? That second question is answerable by a sales manager with no data science, and answering it edits the artefact rather than producing a complaint about it. An opaque score offers no equivalent move: the only available responses are compliance and quiet non-compliance.
Because every dimension emits a written rationale, checking becomes cheaper than redoing — and that ratio is what makes the volume tractable at all. Reading a paragraph and disagreeing with it costs a minute; reconstructing why an account deserves attention costs an hour. While review stays an order of magnitude cheaper than the original work, one person’s judgment can cover four thousand accounts. The number is a sort key. The rationale is the product.
Because the weights were tuned on a test territory first, the blast radius was bounded during the window when the thesis was most likely to be wrong. This is staged rollout — an old engineering idea — applied to judgment rather than to code. A rubric wrong across one territory costs an afternoon. The same rubric wrong across the whole book seeds every downstream conversation with a bad reason, and bad reasons are much harder to un-ring than bad numbers.
And because the tuning happened where the thesis-owner could operate himself, the hand on the dial belonged to the person with the market view. Put a terminal in between and you have inserted a translator into the one loop that has to stay tight.
The transferable principle is one line: encode the judgment, not the answer. The counterfactual is the same book scored by a fitted model — a ranking arrives, nobody can interrogate it, and the interrogation was the work.
Where it breaks
The post does not discuss failure modes; this is our assessment.
- The rubric encodes one person’s theory of the market, then applies it at full confidence to every account. Intrinsic to the practice — and the rationale makes it worse, not better. A wrong ranking with an articulate explanation is harder to dislodge than one with a bare number. Symptom: every rationale reads well and none surprises you.
- The tuning loop has no ground truth in it. Nudging a weight down because it feels heavy calibrates against the author’s expectation, not closed-won. Nothing in the published account closes that loop. Symptom: a year later nobody can say whether the weights were right.
- Proxy dimensions decay. Counting AI mentions on a jobs page discriminates only while posting AI roles is a choice. Symptom: a dimension where nearly every account scores high, consuming weight and adding no ordering.
- A snapshot ages. Scores from one overnight run reflect that night’s research, and the post reports no refresh between the annual planning runs. Symptom: a rep opening a call with a rationale two quarters stale.
- Segment rubrics multiply. Two is a morning of maintenance a year. Eleven is somebody’s second job.
Present the output as a prioritisation aid, not a validated model. And don’t do this at all if you cannot name your dimensions before you open the tool: ask a model to propose them and you adopt a theory of your market that nobody chose, and never notice.
What has to be true at your company
Somebody owns the market thesis and will write it down. This practice does not generate a point of view; it scales one. If no named person will put five dimensions on paper and be wrong about them in public, you get a rubric nobody defends.
The book is big enough that ranking is a real problem. Below a few hundred accounts a team holds the whole book in working memory, and the rubric costs more than the ordering is worth.
Account-level data exists, is queryable, and is fresh enough to score on. Bryant’s run used web research plus Salesforce and BigQuery; the shape that matters is one system of record for pipeline and one for usage or spend. Dimensions keyed to a field reps update sporadically will produce confident sentences built on stale data.
Someone owns territory assignment and will act on the output. A ranking that changes neither territory boundaries nor sequencing changes nothing.
The culture tolerates a leader publishing an opinionated ranking that reps will contest — and has a forum where contesting it lands somewhere.
Across company shapes: at eight people, run one rubric over two hundred accounts in a spreadsheet, founder writing the dimensions — the value is forcing the thesis into words. At eight hundred the risk inverts: the rubric stops being a working document and becomes policy, so version it, date it, name its owner. An unowned rubric outlives the thesis it encoded and keeps ranking the book against a market that has moved.
Try it this week
Take one rep’s territory — fifty accounts is plenty. Before opening any tool, write five dimensions on paper with a weight for each, summing to 100. Twenty minutes, one person, no approval needed. Then score ten accounts against it by hand, writing one sentence of rationale per dimension. Another thirty minutes. That is the whole first move: no automation yet.
What you should see if it is working: at least two of the ten move materially up or down against the current gut ranking, and you can name the dimension that moved them. Both halves matter — movement without a nameable cause is noise.
How you would know it is not: the ten come out in roughly the order you would have written from memory. Then the rubric is restating your priors rather than testing them, and automating it will scale a ranking you already had. Fix the dimensions or stop — do not scale it and hope.
Only once ten hand-scored accounts have told you something you did not already believe is it worth automating the rest of the territory.