← Small Fish

What this actually gets right, and how we know.

Every figure here is measured against businesses a person labelled by hand, without seeing what the engine said. None of it is generated from this page — it is read from the same table the plan is run off, so a number here cannot be newer or older than the number we work from.

measured between 2026-09-20 and 2026-09-23

The four numbers that matter

Precision on matches we charge for
100.0% ✓ PASSES — 41 positive calls, 0 false positives, CI 91.4–100%. 72 enriched hand labels. Dental only; the gate wants three niches
target ≥ 90% · 2026-09-22
Recall against a hand-built true-match set
55.0% (11 of 20), CI 34.2–74.2% — UNDECIDED, no longer failing. Was 30.0% before generic detector hits were escalated to the model
target ≥ 60% · 2026-09-22
Couldn't-tell, on sites we could read
24.5% ✓ over 1,000 businesses — thin margin, 11.8% if generic hits settle
target ≤ 25% · 2026-09-22
What a credit costs us, by band
band 1 $0.0294 · band 2 $0.0687 · band 3 $0.0596 ✓ — replaces "blended cost per match ≤ $0.04", which priced against a threshold no buyer had validated
target ≤ $0.07 · 2026-09-23

And the four that should be read next to them

A benchmark page that only lists what went well is marketing wearing a lab coat. These are the measurements that limit the ones above, and they are here rather than in a footnote.

Run-to-run noise: the same code, the same corpus, twice
precision read 71.4% then 83.3% on two runs that differed in nothing. Below the Haiku→Sonnet gap, no engine change is distinguishable from this
no target — tracked, not gated · 2026-09-21
Businesses whose listed website is wrong — a ceiling on precision
5 of 70 (7%) — caps achievable precision
no target — tracked, not gated · 2026-09-21
The ceiling the crawl imposes on recall
75% (15 of 20 true matches readable). At 55.0% delivered, 3 abstentions separate the engine from the 60% target
no target — tracked, not gated · 2026-09-22
Couldn't-tell, all causes separated
19.0–25.3% ✓
target ≤ 25% · 2026-09-20

Precision is proven on one niche of three. The labelled set is dental. Med spa and HVAC are not labelled yet, so the figure above is not a cross-niche claim and is not used as one. The gate in our own plan wants all three before we buy traffic against an accuracy claim.

Two identical runs disagreed with each other. Same code, same frozen corpus, same labels — and precision came back different, by the margin in the noise-floor row above. That spread is wider than the gap between the models we tested, which means no engine change smaller than it can honestly be called an improvement. It is why every rate on this page is an interval.

Some source records point at the wrong website. The share is in the row above, and it is a hard ceiling on precision that no amount of better reading fixes. It is ours to carry, not the business’s.

The rest of the measurements

Precision on absence criteria alone
100.0% ✓ (lower bound 91.4%, 41 calls) — the gate reads this separately and it passes too
target ≥ 90% · 2026-09-22
Proofs found verbatim in the fetched page
100% (11 of 11 model verdicts)
target high · 2026-09-21
Cost to read and judge one business, cold
$0.0032 Sonnet 5 ✓ over 1,000, 696 cold-fetched (Haiku $0.0011, Opus $0.0074)
target ≤ $0.010 · 2026-09-21
Open-data coverage against Google
veterinary 73.1–86.0% ✓ · dental 63.8–94.3% ? · med spa 45.4–79.3% ? · HVAC 31.1–47.8% ✗
target ≥ 70% · 2026-09-21
Criteria settled with no model call at all
35 of 100 dental criteria, live. 56.0–57.4% where a detector exists (2026-09-20) still stands
no target — tracked, not gated · 2026-09-21

Targets with no number yet

Published so that the list above cannot read as complete. Each of these has a target in the plan and no measurement behind it.

How it is measured

  1. 1 · A market is drawn from open data, deduplicated, and read with a polite crawler that honours robots.txt and identifies itself.
  2. 2 · A sample is labelled by hand with the engine’s verdict withheld. A labeller who can see “the engine said match” agrees with it more often, which inflates the very number the benchmark exists to test.
  3. 3 · Every match carries a quote, and the quote is checked against the fetched text character by character. A verdict whose proof is not found verbatim is a failure whether or not the verdict was right.
  4. 4 · Our own failures — proxy errors, timeouts, network faults — are excluded from every rate rather than counted against the business. A site we could not reach is our problem, not evidence about them.
  5. 5 · Rates are reported as intervals. A sample of 41 does not support a point estimate, and the noise floor above is why.

When a measurement contradicts something we said earlier, the retraction is written down next to the original. There are several.

See what it can find