← Back to Blog
AI Citation VisibilityGEOMeasurement

AI Share of Voice Is a Ratio Built From a Sample. Here's the Arithmetic Nobody Shows You.

2026-09-01
AI Share of Voice Is a Ratio Built From a Sample. Here's the Arithmetic Nobody Shows You.
Contents
How is AI share of voice actually calculated?What does the check budget do to the number?How much of the movement is the tool rather than the site?What is the tool structurally unable to see?What do I track instead?How should you read a share-of-voice report you're handed?

Every AI visibility dashboard leads with share of voice: of all the AI answers about your topic, what percentage mention you. It's the right instinct — a single comparable number, competitor benchmarking built in.

It is also a fraction where both the numerator and the denominator are estimates produced by the vendor, and almost nobody selling you the number will walk you through how it's assembled. I've been tracking it on client sites for months, so here's the walkthrough.

How is AI share of voice actually calculated?

There is no index of AI answers to query. Nobody can see what ChatGPT told ten million people yesterday. So every tool in this category does one of two things:

Prompt monitoring. The tool holds a list of prompts, sends them to each engine on a schedule, reads the answers, and counts mentions. Your share of voice is the share of those answers, on that prompt list, on that schedule.

Dataset indexing. The tool collects a large pool of AI conversations by other means and lets you explore how brands appear across it. Broader, but you don't control the pool and it's a sample of somebody's traffic, not yours.

Either way the denominator is a sample. That is not a flaw to be fixed — it's the only thing possible. The flaw is presenting the output as a measurement of reality.

What does the check budget do to the number?

This is the arithmetic that matters and it's usually buried in a pricing page.

Prompt-monitoring tools sell you checks, and the standard definition is: one check = one prompt × one engine × one location. On one widely used plan, the allowance is 2,500 checks a month.

Watch what happens when you spend it. Say you track 100 prompts across 6 engines:

Setup Checks per run Runs available per month
100 prompts × 6 engines × 1 location 600 ~4
100 prompts × 6 engines × 3 locations 1,800 ~1
30 prompts × 6 engines × 1 location 180 ~13

At four runs a month, each engine-location cell for a given prompt is being sampled once a week. Your "share of voice moved from 18% to 24% this month" is a comparison of two weekly snapshots of a 100-prompt list. For a local service business that wants city-level truth, the location multiplier eats the budget outright.

Anyone who has run A/B tests will recognise the situation: a small sample, a noisy instrument, and a percentage change being read as a trend.

How much of the movement is the tool rather than the site?

I have a clean natural experiment, and it's the reason I stopped treating these numbers as measurements.

Tracking one engine got paused in my panel in mid-July, which means its figure was frozen — no new data collection at all. Over the following weeks that frozen number read 194, then 186, then 163, then 60, then 25.

A metric that cannot receive new data fell eightfold. The only possible explanation is that the vendor's index behind it was being recomputed retroactively. Which tells you what to do with a 6% month-over-month move on a live engine: nothing.

The same series had ChatGPT going from 824 to 813 between two snapshots. I could have written that up as a decline. Pages-cited was unchanged at 22 across the same two readings, so it was sample churn. A 1.3% move in this instrument is not a result.

None of this means the tools are bad — I keep paying for one, and the underlying trend in that series was real and large: 63 responses to 2,605 across six weeks, with the full snapshot-by-snapshot series published here, corrections included. Large moves over long windows survive the noise. Small moves over short windows are the noise.

What is the tool structurally unable to see?

Two things, and both are load-bearing.

The engines it doesn't cover. In the panel I was using, Claude wasn't a tracked engine at all. In my own server log over the same period, Claude-User — the agent that fetches a page while composing an answer for a person — hit one client site 129 times in eight days. That presence existed, mattered, and would have been reported to a client as zero. A share-of-voice denominator built on six engines silently asserts that the seventh doesn't exist.

Whether anyone clicked, called, or bought. Presence is upstream of everything that pays. On the same client, six weeks of thousands of AI answers coincided with 26 recorded visits from all AI platforms combined, because the answer — and in home services the phone number — is delivered inside the chat. Share of voice cannot see that, and a rising share of voice is compatible with zero new business.

What do I track instead?

Not instead of share of voice — alongside it, with share of voice demoted to context rather than headline. Three numbers, in order of how fast they respond to work:

1. Live fetches per URL, from my own server log. The count of times an assistant pulled a specific page while answering somebody. It's mine, it's a census rather than a sample, and it moves within days of rewriting a page. This is the working feedback loop; how it's built and what it can't prove is on the Citation Tracker page.

2. Pages cited, not responses. In that six-week series, responses kept climbing while pages-cited went 13 → 14 → 17 → 22 and then stopped dead at 22. Responses were saturating against a fixed set of citable pages. The number only moves again when new pages enter the set — so pages-cited tells you whether you're growing your citable surface, while responses mostly tell you how the sampler feels this week. Which pages make it in, measured from logs, is the subject of its own post.

3. Leads carrying an origin. The only number that closes the loop, and the hardest to build: a first touch that survives return visits and brand searches into the lead, then into the job and the invoice amount. That's full-cycle attribution, and it's engineering, not reporting.

How should you read a share-of-voice report you're handed?

Four questions. They're the ones I'd want asked about my own reports.

"Which engines are in the denominator?" If Claude, or Copilot, or AI Mode is missing, say so out loud. The excluded engines are being reported as absent. This is not hypothetical: on standard plans one major tool tracks three engines out of six and puts Claude behind Enterprise, which I take apart with log data on peec ai alternative and profound alternative.

"How many prompts, how often, how many locations?" Multiply them. If the product is close to the check allowance, you're looking at roughly one sample per cell per period, and month-over-month moves under about ten percent aren't distinguishable from noise.

"Who wrote the prompt list?" It is the denominator. A list weighted toward phrasings you already win on produces a flattering number honestly.

"What would a decline look like?" If nobody can describe how the instrument would show you losing, it isn't measuring.

Share of voice is a legitimate metric with an illegitimate reputation for precision. Use it for direction over quarters and for competitive comparison at a coarse grain. Don't run a content programme off its weekly wiggle — run it off pages-cited and your own log, which is the practical difference between the three labels everyone argues about and work you can verify.


I run all of it: prompt-level presence tracking, server-side live-fetch logging that no third-party tool provides, and the attribution chain from citation to invoice. What's included and what it costs is public on pricing. If you want a second opinion on an AI visibility report you're already paying for, let's talk.


Related:

Related service pages
AI citation visibilityFrom near-zero to cited across all 7 AI engines in 13 days.