← Back to Blog
AI Citation VisibilityGEOToolsMeasurement

Every AI Visibility Tool Samples Prompts From the Outside. Not One of Them Can See Your Server.

2026-09-01 · updated 2026-09-04
Every AI Visibility Tool Samples Prompts From the Outside. Not One of Them Can See Your Server.
Contents
What do these tools actually do?What does the check budget do to your dashboard?How much of the movement is the tool rather than your site?What can none of these tools see?Then what's the honest weakness of my own instrument?How should you actually buy in this category?

There are now dozens of AI visibility tools, they all show a confident percentage on a dark dashboard, and the category has a structural blind spot that no amount of product polish can close.

I'm writing this from both sides: I pay for one of these tools and use its data in client work, and I built the server-side instrument none of them provide. So this isn't a takedown. It's the thing I wish someone had handed me eight months ago — what the category can see, what it can't, and how to read the number you're being shown.

What do these tools actually do?

Underneath the branding there are only two mechanisms, and one distinction that matters more than any feature list.

Prompt monitoring. The tool holds a list of prompts, sends them to each engine on a schedule, reads the answers, records whether you were mentioned or cited. You control the prompt list, which means you control the denominator.

Dataset indexing. The tool assembles a large pool of AI answers by other means and lets you explore how brands and sources appear across it. Broader than any prompt list you'd write, but it's a sample of someone else's traffic and you don't control what's in it.

The distinction that trips people up: AI search monitoring tracks what users actually see in AI search products, citations and links included. LLM monitoring probes the model's baked-in knowledge and returns no citation data at all. Both get sold as "AI visibility". Only the first can tell you whether your pages are being used.

As published at the time of writing: entry tools start around $29/month (Otterly.AI, which also reports AI crawler analytics and geographic tracking). Peec AI runs roughly $95 / $245 / $495 a month (Starter / Pro / Advanced), with 50 tracked prompts on Starter and 150 on Pro. Profound publishes $99/month Starter and $399/month Growth, both billed annually, with Enterprise on custom terms. Ahrefs Brand Radar puts custom prompt tracking on its all-indexes plan at about $699/month with 2,500 checks, leaning on exploration of a large pre-collected answer dataset rather than only your prompt list. Verify all of it before you buy; this market reprices constantly and these figures already moved once while I was writing.

One limit matters more than any price. Peec's standard plans track three engines at once out of six, with extra models as paid add-ons, and Claude is Enterprise-only. On one client site my log recorded Claude fetching pages 129 times in eight days while answering real people — second only to ChatGPT. A business on Starter or Pro would have been shown zero for that engine. I compare both tools against the server-side view in detail on profound alternative and peec ai alternative.

What does the check budget do to your dashboard?

This is the arithmetic that decides whether your numbers mean anything, and it's the part nobody puts on the landing page.

The industry-standard unit: one check = one prompt × one engine × one location. So a 2,500-check monthly allowance goes fast:

Setup Checks per run Runs per month
100 prompts × 6 engines × 1 location 600 ~4
100 prompts × 6 engines × 3 locations 1,800 ~1
30 prompts × 6 engines × 1 location 180 ~13

At four runs a month, every engine-location cell is sampled about once a week. That is the true resolution of "our AI visibility went from 18% to 24%". For a local business wanting city-level answers, the location multiplier consumes the budget before you've covered your services.

The practical rule: fewer prompts, sampled more often, beats a long list sampled rarely. Thirty well-chosen prompts checked thirteen times a month will tell you more than a hundred checked four times.

How much of the movement is the tool rather than your site?

I have a clean natural experiment on this, and it permanently changed how I read these dashboards.

Tracking for one engine got paused in my panel in mid-July, meaning its figure was frozen — receiving no new data whatsoever. Over the following weeks that frozen number read 194, then 186, then 163, then 60, then 25.

A metric that could not receive new data fell eightfold. The vendor was recomputing its index retroactively. Whatever else that tells you, it tells you what a 6% month-over-month move on a live engine is worth.

In the same series, one engine went from 824 responses to 813 between two snapshots while pages-cited stayed flat at 22. I could have reported a decline. It was sample churn. Single-digit percentage moves in this instrument are not results — the full snapshot series with every correction is published here.

The underlying trend in that same data was real and large: 63 responses to 2,605 across six weeks. Big moves over long windows survive the noise. That's what the category is good for.

What can none of these tools see?

Three things. The first two are the reason I stopped treating any dashboard as complete.

The engines they don't cover. In the panel I was using, Claude wasn't a tracked engine. In my own server log over the same period, Claude-User — the agent that fetches a page while composing an answer for a person — hit one client site 129 times in eight days. That presence was real and would have been reported to the client as zero. Every share-of-voice denominator silently asserts that the engines outside it don't exist. Before you buy, ask which engines are in, and treat the missing ones as unmeasured rather than empty.

Your own server. This is the structural one. These tools observe from outside: they ask engines questions and read answers. They cannot see the request an assistant makes to your site while composing a reply — because that request only exists in your log. On the client site above, that's 425 live answer fetches in eight days across three assistants, against 23 sessions that Google Analytics recorded from all AI platforms in the calendar month of August. No external tool can produce either half of that comparison. Mine can, which is why I built it; what it records and what it explicitly cannot prove is on the Citation Tracker page.

Whether your own server is being read. The counts that answer that are published as an open dataset — per agent, per day, and per URL type.

Whether any of it produced money. Presence is upstream of revenue and the distance between them is where most AI-SEO stories quietly end. No prompt sampler can know a phone rang.

Then what's the honest weakness of my own instrument?

Since I'm auditing everyone else's, mine gets the same treatment.

It classifies by user-agent string, and a user-agent string is a claim. Sorting one site's live-answer fetches, I found requests for /wp-config.php.bak, /wp-config.php.old and /wp/.env arriving with user agents claiming to be assistants fetching pages for people. No assistant asks for a config backup. Someone is forging the header — a scanner using an AI agent name because it's usually allowed through. Proper verification means checking source IPs against the ranges the AI companies publish. Mine doesn't do that yet, so the correct word for my numbers is claimed live fetches. The volume is a rounding error against 425, and it changes no conclusion, but a caveat you drop is how a plausible number becomes a wrong one.

A fetch is not a citation. The agent pulled the page; whether it ended up quoted is a separate question that only a prompt sampler can answer. Which is the actual argument of this post: the two instruments answer different questions and neither replaces the other.

And it cannot see static files. I have zero recorded requests for llms.txt across two sites and 6,676 AI requests — a great contrarian headline that I threw out, because llms.txt is served as a static file and my logger runs inside PHP. The zero was guaranteed before the window opened. That whole trap is written up here.

How should you actually buy in this category?

Five questions, in the order I'd ask them:

  1. "Which engines, and which are missing?" Missing engines are unmeasured, not absent. Get the list in writing.
  2. "What's the check allowance, and what's my prompt × engine × location product?" Divide. That's your real sampling rate. If it lands near one sample per period, ignore small movements entirely.
  3. "Citations or just mentions?" An LLM-knowledge probe with no citation data can't tell you which of your pages is working, so it can't guide content.
  4. "Does it tell me what to change?" Presence data alone does not. The structural work it should be feeding is answer engine optimization.
  5. "Who writes the prompt list?" It's the denominator. A list weighted toward phrasings you already win produces a flattering number honestly.
  6. "What would a decline look like in this tool?" If nobody can describe that, it isn't a measurement.

And the recommendation I'd actually give: buy the cheapest tool that covers your engines and gives citations, then spend the difference on server-side logging. The dashboard tells you where you stand against competitors, in coarse grain, over quarters. Your own log tells you which specific URL an assistant reached for yesterday — a census rather than a sample, and a feedback loop measured in days. One is a report. The other is how the work gets steered.

If you want the frame underneath all this — the four different numbers people mean by "AI visibility" and which instrument sees each — that's in its own post, and the share-of-voice metric specifically gets taken apart here.

Where these subscriptions sit against audits and retainers, priced side by side, is what AI visibility costs. If the quote in front of you has five figures on it and no price on the vendor's website, that is a different purchase with different arithmetic — what separates a $10,000 contract from a $70,000 one. And if the decision is turning into a hire rather than a subscription, the questions to ask an AI SEO specialist are the same ones above, pointed at a person.


I run both instruments for clients: prompt-level presence tracking and the server-side logging the category doesn't provide, plus the attribution chain that connects a citation to an invoice. Scope and prices are public on pricing. If you're paying for an AI visibility tool and want a second opinion on what its numbers mean, let's talk.


Related:

Related service pages
AI citation visibilityFrom near-zero to cited across all 7 AI engines in 13 days.