A 70% match rate sounds like a win until you find out a third of those matches put the wrong person at the wrong company. False positives in website visitor identification data don't just waste a name in a spreadsheet, they burn SDR trust, pollute your CRM, and get an entire tool written off as "junk data" after one bad week. If you've bought a visitor identification tool and never audited its output for false positives, you're flying on faith.
This is the audit we run internally and recommend to every customer: a repeatable, quarterly check that catches drift before it corrupts a quarter of pipeline data, not a one-time vendor bake-off before you sign a contract.
What a False Positive Actually Looks Like in Visitor ID Data
A false positive isn't always obvious. It rarely shows up as a blank record, it shows up as a confident, fully-populated one that's simply wrong. The three most common patterns:
- Wrong company, right traffic. The visit was real, but the resolved company has no plausible connection to it, a hospital system identified from a visit to a dev-tools pricing page, for instance.
- Wrong person, right company. The company match is defensible, but the named contact hasn't worked there in a year, or belongs to a department nowhere near the buying motion.
- Stale identity, right graph. The match was correct when the identity graph last refreshed, but the person changed jobs and the record never caught up.
All three look identical in a CRM export: a full name, a title, an email, a company. That's exactly why volume alone can't catch them. You have to go looking.
Why This Matters More Than Your Headline Match Rate
Every vendor, including us, publishes a headline identification number. Ours is 93%* at the account and company level and 62%* at the person level, both measured against engaged sessions (a visit lasting 10 seconds or longer, or two or more pageviews). Results vary by traffic profile, geography, and industry. That number tells you how much of your traffic gets resolved. It tells you nothing about whether the resolution was correct.
A 50% match rate with 90% precision beats an 80% match rate where a third of the matches are wrong. The lower number gives your SDRs a smaller list they can trust. The higher one gives them a bigger list they'll eventually stop believing, and once a rep decides the tool is unreliable, they stop working the good matches along with the bad ones. That's the real cost of an unaudited false-positive rate: not the wasted records themselves, but the rep who quietly disengages from the entire channel.
If you haven't stress-tested a vendor's accuracy claim before buying, our guide to vetting a vendor's accuracy claims covers the pre-purchase bake-off. This post picks up where that one leaves off: the ongoing discipline once the tool is already live and feeding your CRM every day.
The Quarterly False-Positive Audit
You don't need a data team or a statistics background to run this. You need one hour, a spreadsheet, and a sample of last quarter's identified records.
- Step 1: Pull a random sample, not your best-looking records. Export 100 identified sessions from the last 90 days at random. Don't cherry-pick the accounts your reps closed, those are survivorship bias, not a quality check.
- Step 2: Cross-reference company matches against firmographic reality. For each record, check industry, employee count, and headquarters location against a source like LinkedIn or the company's own site. A B2B SaaS company matched to a regional hospital network is a false positive regardless of how confident the record looks.
- Step 3: Verify person-level matches against current employment. Spot-check named contacts on LinkedIn. A title and company that were accurate eight months ago but aren't today counts as a false positive for this quarter's purposes, even if the match was technically correct once.
- Step 4: Score by severity, not just by count. Use a tiered scale so the audit produces a decision, not just a number:
- 🔴 Very High - Wrong company entirely, no plausible connection to the visiting traffic. Escalate to your vendor immediately.
- 🟠 High - Right company, person no longer employed there or in an unrelated function. Route back into enrichment, don't route to an SDR.
- 🟡 Medium - Right company and person, but title or seniority has drifted since the identity graph last updated. Usable with a manual title check.
- 🟢 Low-Medium - Minor formatting or normalization issues (a legal entity name instead of a brand name) that don't affect targeting.
- ⚪ Low - Clean match, verified current, ready to route as-is.
Step 5: Calculate your working accuracy rate. Divide records scoring 🟢 or ⚪ by your total sample. That number, not the vendor's published match rate, is what should govern how much you trust auto-routed records versus records that get a human glance first.
Building the SDR Feedback Loop
An audit that happens once a quarter and changes nothing is a compliance exercise, not a quality process. The audit only pays off if bad matches get routed somewhere useful instead of silently sitting in your CRM as noise:
- Give reps a one-click flag, not a support ticket. If flagging a bad match takes more than five seconds, reps won't do it, and you lose your best real-time signal between quarterly audits.
- Route flagged records to a suppression list, not the trash. A pattern of false positives from one account, one industry, or one traffic source is diagnostic information. Deleting the record throws away the diagnosis.
- Watch for clustering, not just individual misses. Ten scattered false positives across ten industries is noise. Ten false positives all from the same referral source or the same ISP block is a fixable identification gap, not random error.
- Close the loop with your vendor on 🔴 and 🟠 patterns. A defensible identification vendor treats a documented false-positive pattern as a data point to fix, not a support ticket to close. If a vendor gets defensive about a well-documented pattern instead of curious, that's a signal on its own.
This is easier to sustain when the audit and the feedback loop live inside the same workflow automation that's already routing identified visitors to your reps, rather than as a separate spreadsheet nobody remembers to open. If bad matches are already flowing through the same play logic as good ones, a suppression rule takes minutes to add, not a quarter to plan.
What a Defensible Identification Layer Looks Like Under Audit
On the mechanism: person-level identification works by matching visitors against an identity graph built from a consent-based publisher network, not by unmasking anonymous individuals from thin air. A graph like that improves the way any graph does, through coverage and cross-validation, which means match quality should trend up across your quarterly audits over time, not stay flat or drift down. If your audit shows the same false-positive rate quarter over quarter, or a rising one, that's worth a direct conversation with your identification provider about what's changing in their graph. Flat or improving accuracy is the bar. Silence when you ask about it is the red flag.
Once your identification data has passed its own audit, the next place to look is how efficiently it's converting once it's trustworthy. Our breakdowns of CRM data hygiene for identified visitors and building a no-noise Slack alert playbook both assume clean input data as a starting point, this audit is how you make sure that assumption holds. And if you want to see exactly how a published match-rate number can be tested against real traffic, our teardown of why no two RB2B match rate numbers agree walks through the same verification instinct applied before you buy.
FAQ
How often should I audit my website visitor identification data?
Quarterly at minimum. Identity graphs update continuously, matching logic changes, and traffic mix shifts, so a false-positive rate that looked fine in Q1 can drift by Q3 without any single dramatic failure.
What sample size is enough to catch false positives?
100 randomly sampled records is enough to get a directionally reliable read for most B2B sites. Smaller sites with lower identified volume may need to pool two or three months of records to hit that number.
Should I stop using a vendor after finding false positives?
Not automatically. Every identification method produces some false positives, the question is whether the rate is disclosed, stable, and improving, and whether the vendor engages seriously when you bring them a documented pattern.
Is a false positive the same as a low match rate?
No, and conflating them is the most common mistake in this category. Match rate measures how much of your traffic gets resolved. False-positive rate measures how much of what got resolved is actually correct. A tool can have a high match rate and a high false-positive rate at the same time.
Who should own this audit, marketing or sales ops?
RevOps or sales ops, since they own CRM data quality and are best positioned to build the suppression-list and feedback-loop mechanics. Marketing should stay looped in when the false-positive pattern affects intent scoring or ABM targeting.
Want your identification data audited against your own traffic instead of a vendor's headline number? Book a Knock2 demo and see what a transparent identity graph looks like under a real audit.




