Lead scoring a salesperson will actually trust
A score gets ignored the first time it disagrees with a rep and cannot say why. What makes one explainable, and why outbound volume must never count.
· 4 min read
Why most scores get ignored within a month
A lead score that a salesperson cannot explain gets treated as noise the first time it disagrees with their own read of a deal, and after that it is ignored permanently, not just that one time. This is not a training problem or a change-management problem; it is a design problem. A score built as a single opaque number — the output of a model nobody in the room can walk back to its inputs — asks for trust it has not earned, because there is nothing to check it against except a gut feeling, and a gut feeling that disagrees with a black box will win almost every time.
The fix is not a more accurate model. It is a score a person can take apart.
Score only what the customer did
The inputs that actually predict anything are things the customer did — replying, moving through a stage a human assessed, having a deal size attached — not things the business did. This distinction matters more than it sounds: a score that rises with outbound message volume rewards exactly the behaviour that gets a WhatsApp number restricted, as our guide to bulk sending and number restriction explains, because five follow-ups nobody answered is not engagement, it is a business talking to itself and calling it progress.
A scoring system built around this rule reads events, not intentions: a reply, a stage change, a deal value being set, a stretch of silence. It has no field for a message the business sent, on purpose, so there is no way to inflate a score by sending more rather than by earning a response.
A reason for every point, not just a total
A useful score shows its work: this lead has a value of 22 because of a reply in the last 90 days worth 12 points, a not-yet-contacted stage worth 0, and a set deal value worth 10. A rep who disagrees with 22 can now disagree with a specific line, which is a conversation worth having, instead of disagreeing with an opaque total, which is not. Built this way, the components typically break down as replies contributing up to 48 points (weighted per reply but capped, since twenty messages from one confused person is not twenty times as promising as one clear one), the pipeline stage a human already assessed contributing up to 40, a set deal value contributing a flat 10, and a staleness penalty for prolonged silence pulling points back down.
That last component matters as much as the positive ones: a lead that was hot three weeks ago and has said nothing since is not still hot, and a score with no way to fall is a score that eventually describes yesterday's pipeline, not today's.
A rolling window, not a decay curve
Replies are counted within a fixed recent window — the last 90 days, adjustable per workspace — and a reply older than that simply drops out of the count. An exponential decay curve is the more mathematically elegant way to handle this and, deliberately, the wrong choice here: when a rep disagrees with a score, 'we counted the four replies from the last 90 days' is a sentence they can check against their own memory of the account. A half-life is not something anyone holds in their head, and a component nobody can sanity-check against their own recollection is a component that erodes trust the first time it produces a surprising number for a reason nobody can articulate on the spot.
Letting the business add its own rules without breaking the trust
The built-in components will not cover everything a workspace cares about — a lead from a specific channel, a title containing a specific product name, silence past a specific threshold. Rather than override the built-in score, a workspace should be able to add its own criteria in a small number of fixed, transparent weight buckets — strongly positive, mildly positive, negative — instead of typing in an arbitrary number that nobody else can calibrate against. A preview that shows exactly what a candidate rule would do to a real, specific lead before it is saved closes the loop: a rule that looks right in the abstract but does something surprising to an actual pipeline gets caught before it goes live, not after a quarter of leads have been mis-prioritised by it.
What this deliberately does not do yet
Arithmetic, not a trained model — and that is a decision, not a limitation to apologise for. A model needs outcome data to learn from, and a business with few or no closed deals in its system has nothing for a model to train on; a model trained on almost nothing produces confident-looking noise, which is worse than a transparent formula, not better. When enough won and lost deals exist to learn from, that data becomes the honest baseline a smarter model would have to beat — not before. And because the score deliberately never rewards outbound volume, it also cannot yet reward a rep's own skill at working a deal beyond what shows up as customer replies and stage progress; that is a real, current limitation, named rather than hidden.
Common questions
Why shouldn't a lead's score go up when we send it more follow-up messages?
Because that rewards effort a business puts in, not interest a customer shows — and a score built to reward outbound volume creates an incentive to send more, which is exactly the sending pattern that risks a WhatsApp number's standing. A trustworthy score only moves on what the customer actually did: replying, being moved to a further stage, having a value attached to the opportunity.
Isn't a rolling window less sophisticated than a proper decay function?
Mathematically, yes. Practically, a rolling window is the version a salesperson can actually verify against their own memory — 'we counted the replies from the last 90 days' is checkable in a way a half-life constant is not. A score's job is to be trusted and acted on, not to be the most elegant function that could have produced a similar number.
How much can a business customise a lead score without breaking the built-in logic?
By adding rules in a small number of fixed, transparent weight buckets rather than overriding the built-in components. The built-in signals — replies, stage, value, staleness — stay in place, and workspace-specific criteria layer on top with a preview step that shows exactly what a candidate rule would do to a real lead before it is saved.
Why arithmetic instead of a machine-learning model for lead scoring?
Because a model needs outcome data — real won and lost deals — to learn from, and a workspace with little or no closed-deal history has nothing to train one on. A model trained on almost no data produces confident-sounding noise, which is worse for trust than a transparent formula. Arithmetic is the honest baseline until there is enough history for a model to genuinely beat it.