Visible2 is an AI visibility consultancy that helps businesses understand how AI systems interpret, trust, and recommend websites in AI-driven search and discovery.

Five stacked translucent panels labelled identity, content, authority, structured data and external validation
Methodology

How we measure AI visibility

Published in full, so you can hold us to it — and so anyone else can adopt it. If more of this industry measured properly, all of it would be worth more.

Structured data cannot rescue an identity that is unclear underneath it.

The problem with asking once

Almost every AI visibility audit on the market works the same way. Somebody asks ChatGPT a handful of questions, screenshots the answers, and puts them in a PDF.

SparkToro tested whether that is reliable. Working with 600 volunteers across roughly 2,961 prompt runs on ChatGPT, Claude and Google AI Overviews, they found there is less than a 1-in-100 chance of getting the same list of brands on two runs of an identical question. The odds of the same brands in the same order are roughly 1 in 1,000.

The practical implication is enough on its own: if repeated runs can produce different brand lists, one response cannot support a claim about overall visibility. The useful metric is appearance frequency across a defined set of buyer questions, with the conditions of each run recorded.

So a single test is not a weak measurement. It is close to no measurement at all.

Our measurement protocol

  • Every question run repeatedly, never once
  • Results reported as a rate with a range, never yes or no
  • Logged-out sessions, memory disabled
  • Assistant and mode recorded, so unlike runs are not mixed together
  • Every run timestamped and retained
  • Re-measured monthly, because source use and model behavior change over time

The raw logs come with every report. You are welcome to check our working, and we would rather you did.

Where this goes wrong in practice

A business owner checks one assistant once, sees the company named, and concludes visibility is strong. Another business checks once, does not appear, and concludes it is invisible. Both conclusions are larger than the evidence supports.

SparkToro's repeated-run research shows why: recommendation lists vary materially even when the prompt is identical. That means a screenshot is an example of an answer, not a measurement of a brand's overall visibility.

The same principle applies to supplier reports. If a report does not disclose how many times each question was tested, which assistant and mode were used, and when the runs occurred, the result is difficult to reproduce or compare over time.

The six rules

Why each one is there

Rule Why
Every question is run repeatedly, never once.A single run tells you almost nothing. A rate across many runs tells you where you actually stand.
Results are reported as a rate with a range, never yes or no.A single observation cannot support a binary claim, and a rate without a range hides how uncertain it is.
All sessions are logged out with memory disabled.Persistent memory biases answers toward brands discussed previously, so a logged-in test measures your own history as much as the model.
The assistant mode is recorded.Retrieval behavior and source use can differ by product mode. Recording the mode keeps unlike runs from being mixed into one result.
Every run is timestamped and retained.Citation behavior shifts fast enough that an undated result is worthless.
Measurement repeats monthly, never once.Source patterns can move quickly. Axios reported that Reddit averaged about 3.83% of citations in ChatGPT Search responses between 18 July and 7 August 2026, illustrating why a historical snapshot should not be treated as a permanent rule.
Straight answers

What we will not claim

Several things widely sold in this industry are not supported by evidence, and we do not sell them.

llms.txt as a visibility lever

Ahrefs analyzed 137,210 domains and found that 97% of published llms.txt files received no requests in May 2026. Google also says llms.txt is not needed for its generative AI search features. We can publish one because it is cheap and may be useful to some tools, but we do not score it as a proven visibility lever.

Special AI schema markup

Google states that no special structured data is required for its generative AI features. Valid structured data can help systems understand page meaning and can support eligible Search features, but there is no separate "AI schema" that guarantees generative visibility.

Page-speed thresholds for AI crawlers

Figures like "under 800 milliseconds" and "a 10-second timeout" circulate widely. When traced, they lead to unattributed assertions in blog posts. No AI company publishes a crawl timeout.

Chunking content for AI parsing

Google's own guidance says its systems handle nuanced multi-topic pages without fragmentation.

Indexing on Bing as a guarantee of ChatGPT visibility

Bing indexing can still be useful for search generally, but OpenAI now documents OAI-SearchBot as the crawler used to make websites eligible for ChatGPT search results. Bing is not a substitute for allowing OpenAI's own search crawler.

We are publishing this because the category has a credibility problem, and the fastest way to fix it is for somebody to show their working. If you are comparing suppliers, ask each of them how many times they run each question. The answer tells you most of what you need to know.

See it applied to your own business

The free scan uses the same technical checks. The Roadmap uses the full protocol above.

Visible2 is not affiliated with similarly named brands in telecom, health, or clinical research. Visible2 is an independent consultancy.