What this actually gets right, and how we know.
Every figure here is measured against businesses a person labelled by hand, without seeing what the engine said. None of it is generated from this page — it is read from the same table the plan is run off, so a number here cannot be newer or older than the number we work from.
measured between 2026-09-20 and 2026-09-23
The four numbers that matter
And the four that should be read next to them
A benchmark page that only lists what went well is marketing wearing a lab coat. These are the measurements that limit the ones above, and they are here rather than in a footnote.
Precision is proven on one niche of three. The labelled set is dental. Med spa and HVAC are not labelled yet, so the figure above is not a cross-niche claim and is not used as one. The gate in our own plan wants all three before we buy traffic against an accuracy claim.
Two identical runs disagreed with each other. Same code, same frozen corpus, same labels — and precision came back different, by the margin in the noise-floor row above. That spread is wider than the gap between the models we tested, which means no engine change smaller than it can honestly be called an improvement. It is why every rate on this page is an interval.
Some source records point at the wrong website. The share is in the row above, and it is a hard ceiling on precision that no amount of better reading fixes. It is ours to carry, not the business’s.
The rest of the measurements
Targets with no number yet
Published so that the list above cannot read as complete. Each of these has a target in the plan and no measurement behind it.
- Warm cost per business
- Weekly profile change rate (drives alert cost)
- p95 time to first match, cold market
How it is measured
- 1 · A market is drawn from open data, deduplicated, and read with a polite crawler that honours robots.txt and identifies itself.
- 2 · A sample is labelled by hand with the engine’s verdict withheld. A labeller who can see “the engine said match” agrees with it more often, which inflates the very number the benchmark exists to test.
- 3 · Every match carries a quote, and the quote is checked against the fetched text character by character. A verdict whose proof is not found verbatim is a failure whether or not the verdict was right.
- 4 · Our own failures — proxy errors, timeouts, network faults — are excluded from every rate rather than counted against the business. A site we could not reach is our problem, not evidence about them.
- 5 · Rates are reported as intervals. A sample of 41 does not support a point estimate, and the noise floor above is why.
When a measurement contradicts something we said earlier, the retraction is written down next to the original. There are several.
See what it can find