Six metrics that measured the past, and the two that didn't
A sales leadership team wanted a weekly scorecard. The CRM could count a dozen things. Almost all of them turned out to be measuring work that happened before the week started.
Client anonymised — a national residential builder, delivered under another company. Names, team structures and record identifiers are withheld or changed. Every figure below is real.
- Reps scored
- 14 sales consultants
- Components rejected
- 6 rejected with evidence
- Baseline compliance
- 15.5% of meetings logged with an outcome
- Cross-validation
- 236 against a 240 hand-built baseline
The brief was a number, the problem was a definition
The ask was straightforward: score fourteen sales consultants weekly, out of ten, on things the CRM already knows. The temptation with a brief like that is to reach for whatever is easiest to query — deal completeness, contact hygiene, pipeline movement — and ship a dashboard that looks authoritative by Friday.
The problem is that most of what a CRM can count is a stock, not a flow. It tells you the state of the database today, which is mostly the accumulated residue of every week before this one. Score a rep on that and you are scoring them for history they inherited, which is both unfair and useless as a management signal.
So each candidate component was tested against live data before it was accepted, and six of them died there.
What the evidence killed
Deal completeness looked obvious until the data showed one specific field was populated on exactly one deal in the entire portal. Scoring it would have produced thirteen zeroes and one hero, measuring nothing but a single person's habit.
New-deal completeness looked stronger — until a query over the last thirty days returned 165 deals created by these reps with zero missing region and zero missing sale type. The 52% gap visible across all time was entirely pre-process legacy. The reps were already at a hundred percent; the metric had no room to move and would have scored everyone identically.
Contact hygiene failed for a subtler reason: lead source was 99.6% complete, but only because three workflows write it automatically. It measured the automation, not the person. Missing-phone was set by the lead source, not the rep.
Deals stuck in stage, deal amount, close date and next-activity were all rejected together, for the same reason: each measures a backlog the rep walked into on Monday morning.
The field that separates people from machines
The component that survived — overdue tasks the rep created and did not close — turned out to hide the sharpest trap in the whole build.
Tasks in this CRM come from three sources: created by a human in the interface, generated by a workflow, and produced by a recurring repeat. The obvious approach is to filter them by subject line, because human tasks and machine tasks tend to read differently.
That does not work, and it fails silently. Recurring repeats inherit the subject of the original human task. A task called "Call the client" is correctly attributed to a person the first time and to the recurrence engine every time after, while reading identically in both cases. Any subject-based rule quietly scores people for work a machine generated.
There is exactly one reliable discriminator: the record's own source field. Filtering on it — rather than on anything human-readable — is the difference between a scorecard that measures behaviour and one that measures the automation running underneath it.
Two components, ranking almost inversely
What shipped was deliberately small: two components, five points each, scored weekly, for the fourteen consultants only.
The first was ramped rather than absolute. Baseline compliance on logging a meeting outcome was 15.5% — sixteen of a hundred and three meetings in a fortnight, with sixty-three still sitting as scheduled and twenty-four blank. A bar set at ninety percent on day one would have scored the entire team at zero and been ignored by the following Tuesday, so the first four weeks top out at fifty percent for full marks, and the bar moves afterwards. A rep with no meetings that week scores full marks, not zero — the metric is about closing the loop, not about volume.
The report behind the second component was cross-checked against a hand-built baseline before anyone saw it: 236 against 240, which is close enough to confirm the definition and different enough to explain. That gap is the kind of thing worth chasing before launch rather than after someone's bonus depends on it.
The most useful signal came last. The two surviving components rank the team almost inversely — the strongest performer on one is near the bottom of the other. That is the evidence they are measuring different behaviours rather than the same behaviour twice, which is the failure mode of most scorecards with more components than this one.
The two numbers that lived outside the system
Two of the metrics leadership cared about were generated outside the CRM entirely, and the reflex answer is a spreadsheet.
Instead they got a small custom object and a weekly entry path, so the figures live beside everything else and can be reported on with the same tools. Number fields were left deliberately blank rather than seeded with target values — a seeded target reads as an achieved target the moment someone opens the report.
The reminder is a recurring task rather than a workflow, because this platform has no scheduled or cron trigger at all. Functionally identical, and worth saying plainly to the client rather than describing it as automation it is not.
Status
Complete and live. Two-component scorecard running weekly across fourteen consultants, with both supporting reports on a shared dashboard and the offline figures captured against a custom object.
Most CRM metrics measure the backlog someone inherited. Working out which ones measure this week is the difference between a scorecard people act on and a dashboard nobody opens.
RevOps Diagnostic Audit