If you're asking which buying signals actually predict closed-won revenue, the honest answer is: probably not the ones on your current scoring model. Most B2B teams build intent and engagement scoring off a list someone found in a blog post or a vendor's default settings, not off their own win-loss history. The fix is a backtest: pull your last 20 to 30 closed-won and closed-lost deals, tag which signals showed up on each one, and measure which signals actually correlate with revenue before you let any of them drive routing, alerts, or SDR prioritization.
This matters because the cost of an unvalidated signal is not neutral, it is negative. Every signal you treat as "high intent" that doesn't actually predict a close burns SDR time, trains reps to ignore alerts, and quietly degrades trust in the whole system. Backtesting is the cheapest fix nobody runs.
Why the 40-Signal Listicle Doesn't Help You
Search "buying signals B2B" and you'll get the same genre of content every time: a list of 15, 40, sometimes 100+ things that could indicate intent, from pricing page visits to job changes to competitor comparison searches. These lists aren't wrong, they're just incomplete. They tell you what a signal is. They almost never tell you whether that signal, in your business, with your deal cycle and your ICP, has ever actually preceded a closed-won deal.
That gap matters more than it sounds. A signal that's genuinely predictive for an enterprise security vendor with a nine-month cycle, like a compliance page visit, may be noise for a $200/month PLG tool with a two-week cycle. Generic listicles can't tell you that. Only your own deal history can.
The Backtest Method
You don't need a data science team or a BI tool to run this. It's a spreadsheet exercise that takes an afternoon.
- Pull your last 20 to 30 closed deals, split roughly evenly between closed-won and closed-lost, from the same segment. Don't mix enterprise and self-serve in one pass.
- Tag which signals were present on each deal before it closed: pricing page visits, buying committee formation (multiple people from the same account visiting in a short window), job changes into a buying role, repeat visits to a specific page, third-party intent spikes, whatever your current model tracks.
- Compute lift per signal. For each signal type, calculate the percentage of won deals where it was present, and the percentage of lost deals where it was present.
- Set a hard threshold. A reasonable starting bar: keep a signal if it showed up in 60% or more of won deals and less than 20% of lost deals. Anything that doesn't clear that bar gets dropped from the scoring model, not deprioritized, dropped.
- Rebuild the model with only validated signals, and re-run the backtest again in a quarter once you have more closed deals to test against.
The threshold numbers aren't sacred, adjust them to your deal volume and risk tolerance. What matters is having a threshold at all, instead of an ever-growing list of "signals that seem important."
What This Looks Like Against Real Deal Data
We hear the same pattern on calls with prospects and customers building out their signal strategy. One reason website intent gets treated with suspicion by sales teams is that it's inherently probabilistic. As one prospect put it on a recent evaluation call, "these are directional signals, you will get some red herrings or some false positives, and that's just the nature of how it's derived." Teams that internalize that upfront, that a signal is a hypothesis to test rather than a fact to act on, are the ones who actually run the backtest instead of skipping it.
The failure mode we see more often is the opposite: teams that never validate anything and instead keep adding inputs. One growth-stage SaaS team described an existing scoring setup with 82 separate columns feeding a single lead score, built up over years by different people bolting on one more field. When asked to rebuild it, their own instinct was to cut it down to four or five inputs, not add more, once they actually looked at what was driving the score versus what was just noise. That instinct is correct, and it's exactly what a backtest gives you the confidence to act on.
The most common unvalidated pattern we see across onboarding calls is teams routing purely on an aggregate lead score, something like "anything under 50 gets filtered out," without ever having checked which of the inputs to that score actually correlate with a closed deal. The score feels rigorous because it's a number. It isn't rigorous until it's been backtested.
Which Signal Types Tend to Backtest Well
Every business's numbers will differ, but across the accounts we work with, signal types tend to sort into a consistent rough order of predictive strength. Treat this as a starting hypothesis for your own backtest, not a substitute for running one:
- 🔴 Multi-person buying committee formation - three or more distinct people from the same account engaging within a tight window. This is consistently the strongest single predictor we see, because it mirrors how B2B deals actually get bought.
- 🟠 Repeat visits to pricing or comparison pages from an already-engaged account - not a cold first touch, but a return visit after an existing conversation has started.
- 🟡 Job changes that move a known contact into a buying-authority role - a strong but lower-frequency signal, more useful for account-level prioritization than day-to-day routing.
- 🟢 Single-page repeat visits with no committee expansion - some predictive value, but weak enough on its own that it should rarely trigger outreach by itself.
- ⚪ One-off anonymous page views with no return visit - the weakest signal category, and the one most listicles rank far too highly relative to how often it actually precedes a close.
If your current model treats a single anonymous pageview the same way it treats a three-person buying committee forming in the same week, that's the first thing your backtest will expose.
Build the Scoring Model After the Backtest, Not Before
The order of operations matters. Most teams build a scoring model first, then hope it works. Backtest first, then build the model on only what survives. This is also where a platform like Knock2's identification layer earns its keep: you can't backtest a signal you can't see, and person-level and account-level identification is what turns an anonymous pageview into a taggable, attributable data point tied to a real account and contact in the first place. Once signals are validated, Knock2's lead scoring lets you build the model around exactly the inputs your backtest confirmed, instead of an 82-column spreadsheet nobody trusts.
This connects directly to detecting buying committee formation, since committee signals consistently backtest as the strongest predictor. It also pairs with how you think about signal decay and expiration windows, because a signal that backtests well at 7 days old may not hold at 30. And once you're combining validated first-party signals with third-party intent data, our guide to blending first-party and third-party intent data picks up exactly where this backtest leaves off, while our walkthrough on validating third-party intent data applies the same backtesting logic to signals you're buying rather than capturing yourself.
Ready to Score Only What's Been Proven?
Book a demo to see how Knock2 turns validated, identified signals into an automatic scoring model, without the 82-column spreadsheet nobody trusts.
FAQ
How many deals do I need to backtest a buying signal?
20 to 30 closed deals, split between won and lost, is enough to spot a real pattern for most SMB and mid-market motions. If your sales cycle is long or your deal volume is low, extend the lookback window rather than lowering the bar, since thin data produces false confidence, not a real answer.
What is a reasonable lift threshold for keeping a signal?
A common starting point is 60% or higher presence in won deals and under 20% presence in lost deals. Adjust based on your own deal volume and how costly a false positive is for your SDR team's time.
Should I backtest first-party and third-party signals differently?
Yes. First-party signals on your own site are generally cleaner to backtest because you control the capture. Third-party intent data should be validated separately before you blend it in, since its accuracy varies significantly by vendor and data source.
How often should I re-run the backtest?
Quarterly is a reasonable cadence for most teams, or any time your ICP, pricing, or sales motion changes meaningfully. A signal that was predictive for a self-serve motion can lose its lift entirely after you move upmarket.
What if none of my current signals pass the backtest?
That is a useful, if uncomfortable, result. It usually means the model was built on assumption rather than data. Start with buying committee formation and repeat high-intent page visits, the two categories that most consistently backtest well, and rebuild from there.




