How do you measure AI visibility?
Short answer
By running a fixed set of customer questions through several AI assistants, more than once each, and recording whether the business is named, how it's described and which sources are cited. Because answers vary from run to run, the result is a rate, such as named in six of ten runs, and not a rank.
I measure AI visibility by running a fixed set of customer questions through several AI assistants, more than once each, and recording whether the business is named, how it’s described and which sources are cited. As far as I know, no assistant gives a business owner a report of how it appears in answers, so I have to take the measurement from the outside, by asking. What comes back is a rate, not a rank.
What gets recorded
For every prompt, in every assistant, I write down the same set of observations.
| Measure | What it tells you |
|---|---|
| Named or not | Whether you’re in the answer at all |
| Order and company | Where you appear, and which competitors appear beside you |
| Description | What the assistant says you do, and whether it’s favorable |
| Accuracy | Whether the phone number, hours, services and service area are right |
| Sources cited | Which pages the answer pointed to, and whether any are yours |
I record being named and being cited separately because they are different results.
Why the result is a rate and not a rank
The same question asked twice can return different companies. Google’s results for a search from one spot are fairly stable from one hour to the next. An assistant’s answer is written fresh each time, so one run is an anecdote. I run each prompt several times and report how often the business appeared. A hypothetical line in a report would read: named in six of ten runs for “who should I call for AC repair in Chula Vista”, with two competitors named more often.
A rate can be compared with last month’s rate. A single screenshot can’t be compared with anything.
The steps, in order
It’s the same method I use for rankings, which I call Scientific SEO.
- Agree the list of prompts. Deciding which prompts a local business should test is a subject of its own.
- Run a baseline before any change goes in.
- Change one thing, such as correcting a set of listings or rewriting one service page.
- Run the same prompts again, with the same wording, about two weeks later.
- Keep the change if the numbers moved, and say so plainly if they didn’t.
The ongoing version of this is AI visibility tracking.
What can’t be measured
It’s better to know the gaps before reading any report. The first is how many people asked. I know of no assistant that shares question volumes with the businesses being asked about, so there’s nothing like the search counts Google provides. The second is what one particular customer saw, since answers can differ with the person’s location and earlier conversation, and with the assistant’s settings.
I also can’t measure why the assistant chose who it chose. The cited sources are evidence and stop short of an explanation.
Calls are the last gap. Your website analytics can show visits that arrived from an assistant’s link, but many people read an answer and phone without clicking, so that figure runs low and I can’t count every call an answer produced.
A note on single scores
If a tool or an agency hands you one “AI visibility score”, ask what is underneath it. The useful questions are which prompts were run, in which assistants, how many times each, and on what dates. A score with no prompt list behind it can’t be checked or repeated. The measurement I describe here’s deliberately plain: a list of questions and a count of runs, with a table of what came back, all of which you could rerun yourself to confirm.