Do not buy lead-scoring software because the demo produces a convincing number next to every contact. Buy it only if the vendor can explain what the score is for, which data creates it, how fit and behavior are separated, what happens when data is missing, how the score enters routing, how users can challenge it, and how the model is reviewed after your market changes. A score that looks precise but cannot be governed is usually worse than a simpler model your sales and operations teams understand.
Use the questions below in a demo, RFP, pilot, or renewal review.
The shortlist table
| Question | Why it matters | Red flag |
|---|---|---|
| What decision will this score trigger? | A score needs an operational use | “Higher is better” with no routing action |
| Which data enters the model? | Inputs determine bias and coverage | Vendor cannot separate first-party vs external data |
| Can fit and engagement be separated? | Good company ≠ active buyer | One opaque total with no components |
| How are missing values handled? | Sparse data is normal | Missing automatically becomes negative |
| Can we explain a score change? | Reps need to trust and challenge it | No history or reason codes |
| How does it write back to CRM? | Bad sync can damage workflows | Unclear overwrite behavior |
| Can thresholds differ by segment? | Products and motions differ | One universal MQL cutoff |
| How is the model recalibrated? | ICP and behavior drift | “AI continuously learns” with no review process |
| What is the fallback if the score is unavailable? | Routing must survive outages | Pipeline stops when vendor is down |
| How is pricing tied to records, events, or users? | Scoring economics can change at scale | Cheap pilot, unclear production cost |
1. “What exact decision should the score change?”
This should be the first question because it exposes vague projects quickly.
Possible answers include:
- route a high-fit, high-engagement inbound lead to an SDR within a defined SLA;
- prioritize a rep’s daily account queue;
- suppress clearly out-of-market contacts from expensive manual outreach;
- choose a nurture path;
- identify records that need human review.
If the buyer cannot name the decision, the vendor cannot prove that the score helps.
A common mistake is to treat lead scoring as a ranking contest: every lead gets 0–100 and the highest number goes first. In practice, different motions may need different models. HubSpot’s current scoring documentation explicitly supports separate fit and engagement scores, as well as combined scores with categories that reflect both dimensions. That is a useful reminder that “good account” and “active buyer” are different questions.
2. “Which data is first-party, which is vendor-supplied, and which is inferred?”
Ask the vendor to show the data lineage for a real score.
A model may use:
- CRM firmographics;
- form submissions;
- website behavior;
- email engagement;
- product usage;
- opportunity history;
- external company data;
- third-party intent;
- job changes;
- inferred attributes.
Each source has different coverage, freshness, permissions, and failure modes.
Do not accept “we use thousands of signals” as an answer. More inputs can increase complexity without improving the decision.
Ask for a field-level map: source, refresh cadence, missing-data behavior, and whether the value is observed or inferred.
3. “What happens when the data is missing, contradictory, or stale?”
This is where demos tend to be too clean.
Real CRMs contain duplicated accounts, old titles, partial forms, shared domains, acquisition records, stale lifecycle stages, and manually edited fields.
The vendor should explain whether missing data:
- contributes zero;
- receives a default;
- reduces confidence;
- excludes the record from scoring;
- triggers enrichment;
- routes to an exception queue.
Also ask what happens when CRM values disagree with external data.
A robust answer distinguishes “unknown” from “negative.” A missing employee count does not automatically mean the company is too small. A missing web event does not prove lack of intent.
4. “Can you separate fit, behavior, intent, and timing?”
A single total score may be convenient for routing, but operators often need the components.
Consider four records:
- excellent ICP fit, no recent activity;
- weak fit, heavy content activity;
- excellent fit, active pricing-page behavior;
- unknown fit, strong third-party research signal.
Those should not necessarily receive the same treatment.
HubSpot currently allows engagement and fit scores to exist separately and can create combined score categories. G2’s lead-scoring buyer guide also distinguishes rules-based and predictive approaches and emphasizes the role of demographic/firmographic and behavioral data.
The procurement question is not “Does your product use AI?” It is “Can we see which kind of evidence is driving action?”
5. “How do users understand why a score changed?”
Sales teams stop trusting scores when the number moves but the reason is invisible.
Ask to see:
- score history;
- top positive and negative factors;
- timestamp of important signals;
- threshold changes;
- model-version changes;
- manual overrides;
- whether a rep can report an obviously wrong score.
If the system uses predictive or AI scoring, ask what explanation is available at the record level and what is only available at model level.
HubSpot’s AI lead-scoring workflow, updated in August 2026, can evaluate historical lifecycle changes to recommend criteria for fit or engagement scores. That capability is useful, but the operational question remains: can the team review, edit, and govern the recommendations before turning the score into automated action?
6. “What exactly is written back to our CRM?”
This is a systems question disguised as a marketing question.
Ask the vendor to list every property it will create or modify:
- total score;
- component scores;
- grade/band;
- confidence;
- last-scored timestamp;
- reason code;
- model version;
- raw intent signal;
- routing status.
Then ask whether writes are one-way or bidirectional and how conflicts are handled.
A score that updates every few minutes can accidentally retrigger workflows. An integration that overwrites a manually curated field can create bigger problems than the score solves.
During the pilot, place writeback into a sandbox or controlled field set before allowing it to touch production routing.
7. “Can thresholds differ by product, segment, source, or motion?”
A universal threshold is attractive because it is simple.
It may also be wrong.
Enterprise inbound, SMB self-serve, event leads, partner referrals, and outbound prospecting can have different base rates and different costs of follow-up.
Ask whether the system supports:
- multiple models;
- segment-specific thresholds;
- separate fit and engagement cutoffs;
- source-specific routing;
- time decay;
- exclusions;
- minimum data requirements.
If it does not, understand the operational compromise before signing.
8. “How will we know the model is getting stale?”
This question separates software demos from operating systems.
Markets change. Product packaging changes. Teams enter new regions. A new acquisition channel appears. The company moves upmarket. Historical “won” data may stop representing the next target market.
G2’s buyer guide notes that lead-scoring formulas require ongoing revision and optimization. That is not a minor maintenance detail; it is a procurement requirement.
Ask the vendor:
- what monitoring is built in;
- how often model performance is reviewed;
- whether past model versions can be compared;
- how threshold changes are audited;
- what volume is needed before recalibration;
- whether a customer can pause or roll back a model.
“AI continuously learns” is not enough. You need a review process.
9. “What is the pilot acceptance test?”
Do not define success as “the integration worked.”
Before the pilot starts, agree on evidence such as:
- coverage: what share of eligible records receive a usable score;
- agreement: whether high-scoring groups actually show better downstream outcomes;
- routing quality: whether sales accepts more of the records prioritized by the model;
- stability: whether small data changes cause unreasonable score swings;
- latency: whether scores arrive in time for the workflow;
- explainability: whether operators can diagnose surprising scores;
- workload: how many exceptions need manual cleanup.
Do not demand a universal conversion lift from the vendor. Your baseline rate, sales process, traffic mix, and sample size matter.
10. “What happens if we stop using you?”
Exit questions belong in the buying process.
Ask how to export:
- scores;
- model configurations;
- reason codes;
- historical score changes;
- raw vendor signals where licensing permits;
- field mappings.
Also ask what happens to API connections, CRM properties, workflows, and stored data after termination.
A good scoring system should improve your operating model, not make the business unable to prioritize leads without one vendor.
11. “How does pricing scale when the pilot becomes real?”
Do not compare only the starting subscription.
Ask whether cost changes with:
- seats;
- scored records;
- enriched records;
- intent events;
- API calls;
- data refresh frequency;
- historical data;
- multiple workspaces or business units;
- premium integrations;
- predictive/AI features.
Model the bill at today’s volume and at a realistic 12–18 month volume. The cheapest pilot can become the most expensive production architecture if every useful signal is metered separately.
The next-step checklist
Before choosing a lead-scoring vendor or partner, make sure you can answer all of these:
- We can name the business decision the score will trigger.
- We know which inputs are first-party, third-party, and inferred.
- Missing and contradictory data have explicit handling rules.
- Fit and engagement can be viewed separately when needed.
- Score changes can be explained and audited.
- CRM writeback fields and workflow side effects are documented.
- Thresholds can reflect our segments and sales motions.
- Recalibration and model-version review have an owner.
- Pilot acceptance criteria are written before launch.
- Pricing has been modeled at production scale.
- Exit/export behavior is documented.
The best vendor is not the one that produces the most impressive score in a demo. It is the one that can fit into a governed revenue process, expose uncertainty, survive messy data, and still help a human team make a better priority decision.
Sources
- HubSpot Knowledge Base, Overview of the lead scoring tool — https://knowledge.hubspot.com/scoring/understand-the-lead-scoring-tool — accessed 2026-10-03
- HubSpot Knowledge Base, Build contact lead scores with AI — https://knowledge.hubspot.com/scoring/build-lead-scores-with-ai — accessed 2026-10-03
- G2, Lead Scoring Software Buyer Guide — https://www.g2.com/categories/lead-scoring/buyer-guide — accessed 2026-10-03
- G2, Best Lead Scoring Software in 2026 — https://learn.g2.com/best-lead-scoring-software — accessed 2026-10-03
Related Reading
- https://salesai.globalsiriusmc.com/articles/data-enrichment-composite-case-conflicting-records-and-exit-rule/
- https://salesai.globalsiriusmc.com/articles/lead-scoring-buying-guide-fit-intent-routing-review/
- https://salesai.globalsiriusmc.com/articles/data-enrichment-vendor-checklist-source-coverage-writeback-governance/