Checking ChatGPT once and screenshotting the answer feels like measurement because the output is concrete. The problem is statistical: one generated response is only one sample from a system that can produce different recommendation lists on repeated runs.
That is the central problem with AI visibility measurement. A single answer can be useful as an example, but it is not enough evidence to support a claim about how visible a brand is overall.
The research is unambiguous
SparkToro, working with 600 volunteers across roughly 2,961 prompt runs on ChatGPT, Claude and Google AI Overviews, found there is less than a 1-in-100 chance of getting the same list of brands on two runs of an identical question. The odds of the same brands in the same order are roughly 1 in 1,000.
That is not a rounding error. It means a single test tells you almost nothing.
SparkToro’s result is enough to change the measurement method on its own. If repeated runs can produce materially different brand lists, visibility should be reported as frequency across a defined prompt set — not as a fixed rank copied from one response.
It misleads in both directions
False confidence. You check, you appear, you conclude you are fine. You may have caught the one run in five where you were named.
False despair. An agency sends a report saying you are invisible. You may in fact appear in a third of answers. Without repeated sampling, nobody knows which.
The second is the more expensive mistake, because it gets sold against.
What a defensible measurement looks like
- Every question is run repeatedly, never once.
- Results are reported as a rate with a range, never as a yes or a no.
- All sessions are logged out with memory disabled — persistent memory biases answers toward brands you have discussed before, so a logged-in test measures your own history as much as the model.
- The assistant mode is recorded. Retrieval behavior and source use can differ by product mode, so unlike modes should not be mixed into one visibility rate.
- Every run is timestamped and retained.
- Measurement repeats monthly. Source patterns move: Axios reported that Reddit averaged about 3.83% of citations in ChatGPT Search responses between 18 July and 7 August 2026. A historical snapshot is not a permanent rule.
One question to ask any supplier
“How many times do you run each question?”
If the answer is once, they are selling you a screenshot. If they cannot answer at all, they have not thought about it. The answer to that single question tells you most of what you need to know about whether the rest of their report means anything.
We publish our full measurement protocol, including the raw logs with every report, so you can check our working. We would rather you did.
