Build lead scoring from what actually predicts conversion in your historical data rather than from assumptions about seniority. Most models over-weight job title and under-weight behaviour, which routes non-engaged senior contacts to sales and teaches reps to ignore the queue.
Why most models fail
The typical scoring model is built in a workshop. Someone suggests that VPs are more valuable than managers, that a demo request is worth more than a whitepaper download, and that companies over 500 employees should score higher. Numbers get assigned by intuition and the model goes live.
It is never validated against whether those signals actually predicted conversion in past deals, and it is rarely revisited. Within two quarters sales has noticed that high-scoring leads convert no better than low-scoring ones and has stopped using the queue.
Build it from history instead
- Pull two years of leads with outcomes. Which became opportunities, which became closed-won, which went nowhere.
- Test each attribute against conversion. Did seniority actually correlate? Did company size? Did visiting the pricing page?
- Rank attributes by predictive strength, not by intuitive appeal. The results usually surprise people.
- Weight accordingly, and discard signals that did not predict anything.
- Validate on a holdout period. Score leads from a period you did not use to build the model and check whether high scores converted better.
- Publish the validation to sales before asking them to trust it.
That last step is the one most often skipped and the one that determines adoption. Sales does not distrust scoring because they dislike models; they distrust it because the last one was wrong and nobody proved this one is different.
Behaviour usually beats demographics
In most B2B datasets, behavioural signals predict conversion better than firmographic ones. Repeat visits to pricing pages, return visits within a short window, engagement with bottom-funnel content and multiple contacts from the same account all tend to outperform job title.
That does not make demographics worthless — an ICP mismatch should be disqualifying regardless of engagement. The useful structure is usually a fit score that acts as a gate and a behaviour score that acts as a ranking, rather than one blended number that obscures both.
Fit and intent as two dimensions
| High fit | Low fit | |
|---|---|---|
| High engagement | Route to sales immediately | Investigate — often a competitor or student |
| Low engagement | Nurture, alert on next engagement | Suppress or exclude |
This two-dimensional view is more actionable than a single score because the four quadrants imply genuinely different actions. A single blended number cannot distinguish an engaged non-buyer from a disengaged perfect-fit account, and treating them the same is how good leads get buried.
Decay and recency
Engagement that happened four months ago should not score the same as engagement from yesterday. Without decay, contacts accumulate points indefinitely and eventually cross the threshold on the strength of long-stale activity, arriving at a rep who finds a conversation from a buying process that ended in spring.
Implement score decay over a window matched to your sales cycle. It is a small configuration change and it removes a category of false positives that damages trust disproportionately, because a stale lead is the most memorable kind of bad lead.
Keeping trust once you have it
Review the model quarterly against actual conversion. Publish the results to sales, including when the model performed poorly. A model that is visibly maintained retains credibility even through a bad quarter; a model nobody has touched in a year loses it permanently at the first visible failure.
Also make rejection easy and structured. When a rep rejects a scored lead, capture the reason in a fixed list rather than free text. Those reasons are the highest-quality training data available for the next model revision, and collecting them signals that the feedback is genuinely wanted.