An AI visibility audit that runs a prompt once and screenshots the answer has measured nothing. AI answers change from run to run, so one missing citation on one run is not enough to establish a visibility problem.
A useful audit samples real buyer questions repeatedly, tracks mentions and citations separately for each engine, explains the likely cause of each gap, and states how every fix will be re-tested.
Below is what each part involves, followed by a sample finding written the way a client would receive it.
Key Takeaways
- One run of one prompt is not evidence. In a Kevin Indig and AirOps analysis, only 2.2% of ChatGPT citations remained after the same prompt was run three times.
- Mention rate and citation rate measure different things. Report both, per engine, and never blend engines into one score.
- Rerunning the same prompt has sharply diminishing returns. One 2026 preprint found that spreading measurement across models and languages reduced error far more efficiently.
- Every finding should keep evidence, observation, and inference visibly separate, with a severity rating attached.
- A fix without a re-test plan is a guess with a deadline.
Why can’t one prompt prove you weren’t cited?
Because the same prompt can return very different sources from one run to the next.
Kevin Indig, working with AirOps, analyzed 815,000 prompt-page pairs and found that after three runs of the same prompt in ChatGPT, only 2.2% of citations remained (Search Engine Land, June 2026).
The drift continues week to week. In the same piece, Indig cites a SISTRIX dataset of 82,619 prompts tracked over 17 weeks: Google AI Mode replaced 56% of its cited sources each week, and ChatGPT replaced as much as 74%.
Smaller datasets show the same pattern. Bill Widmer ran 72 buyer questions through four models two days apart; Claude kept 30% of the URLs it had cited, and Perplexity, the most stable, kept 65% (Orbit Media, September 2026).
So “we asked ChatGPT and you weren’t there” is a starting question, not a conclusion. The audit’s first job is designing a measurement that can survive that much variance.
How should an audit sample AI answers?
With a fixed set of real buyer questions, asked in several wordings, several times, on each engine separately. Each of those choices removes a different kind of noise.
The question set. Questions should come from how buyers actually talk: sales calls, enquiry emails, support messages. Widmer recommends exactly this over keyword tools.
Group them by intent: questions naming the brand, questions about the category, and questions describing the problem. A business can be visible for one group and absent from another.
Engines tracked separately. ChatGPT, Perplexity, Gemini, Google AI Overviews and Claude cite different sources.
In Widmer’s data, all four models tracked agreed on citing the same domain for the same question in 1.7% of cases. Indig compares blending them into one score to averaging your Google rank with your Bing rank.
Breadth, not just repeats. A July 2026 preprint by Dmitrij Żatuchin decomposed 12,933 AI responses about 20 brands and found that spreading measurement across models and languages reduced error far more efficiently than simply repeating the same prompt. Additional repeats produced sharply diminishing returns (arXiv 2607.13304).
That study has real limits. It measured the sentiment of each response toward a brand, not citation stability or citation likelihood, so it does not show that rewording improves citation measurement specifically.
It also used one corpus of Central and Eastern European brands, is not yet peer reviewed, and its author sells AI visibility measurement.
A workable starting design for a small business is 15 to 20 questions, three wordings each, three runs per wording, per engine.
That is a working rule, not an industry standard. Even at that size, the results are directional, and the report should say so.
What is the difference between a mention and a citation?
A mention is when the AI names your business in its answer. A citation is when it links your page as a source.
An audit needs both numbers, because they move independently.
The gap between them varies by engine. In Widmer’s dataset, Claude mentioned the tracked brands in 70% of answers but cited them in 40%, while ChatGPT mentioned them in 56% and cited them in 47%.
The gap also varies by brand size. Semrush was mentioned in 83% to 100% of relevant answers on every model, yet its citation rate on those answers ranged from 83% down to 0%. Widmer’s two small sites showed gaps of only 4 to 5 points.
His explanation is that models may recommend famous brands from training data while citing third-party pages as the source. That is a hypothesis, not a finding.
For a small business, the practical reason to track both numbers is that a mention and a citation can tell very different stories.
Where does an audit look for the cause?
In four places, checked in order, because each one can block the next.
There is no point rewriting a page that AI crawlers cannot reach.
Access. Can AI crawlers fetch the page at all?
The audit checks robots.txt rules, firewall or bot-protection settings, and noindex or nosnippet directives in page code and HTTP headers. A single blocked crawler can explain an entire engine’s silence.
Extraction. Does the page answer the question early and plainly?
Growth Memo’s analysis of 18,012 verified ChatGPT citations found that 44.2% came from the first 30% of the page (Growth Memo, February 2026).
That finding is scoped to ChatGPT, but it gives the audit a concrete test: read the first 100 to 200 words of each key page and ask whether they answer the buyer’s question.
Entity clarity. Can a machine tell who the business is, where it operates, and which profiles belong to it?
The audit checks that the name, services, and location match across the website, LinkedIn, and directories, and that structured data connects them:
{
"@context": "https://schema.org",
"@type": "ProfessionalService",
"name": "Example Coaching",
"url": "https://www.example.com",
"areaServed": "Portugal",
"sameAs": [
"https://www.linkedin.com/company/example-coaching",
"https://www.exampledirectory.com/listing/example-coaching"
]
}The sameAs property explicitly links the organization to pages representing the same entity. That can reduce ambiguity about which profiles belong to the business.
No AI lab has published documentation confirming that schema directly influences which sources get cited, so the audit treats it as a clarity signal, not a citation lever.
Third-party sources. Which pages get cited instead of you?
Widmer found that models mostly cited niche listicles, review platforms, and comparison posts, and he recommends counting which of those repeat and pitching them.
Google rankings are a weak proxy here: in his data, only 10% to 30% of cited domains also ranked in Google’s top 10 for the matching query.
What does a finished audit finding look like?
Here is one finding from an illustrative audit. The business is invented and every number is made up to show the format. None of it is client data.
The business: Harbour Leadership, a leadership coach based in Porto who works with first-time managers.
The severity scale used in the report:
- Critical: AI systems cannot access the site or a key page.
- High: absent from a core buying question on an engine the audience uses.
- Medium: present but uncited, inaccurate, or inconsistent.
- Low: hygiene issues with no observed effect yet.
Finding 03: Not cited for the core service question on Perplexity
Severity: High
Evidence: Three wordings of “Who is a good leadership coach for first-time managers in Portugal?”, run three times each on Perplexity during one week in September. Harbour Leadership was mentioned in 1 of 9 answers and cited in 0 of 9.
Two coaching directories were cited in 7 of 9 and 6 of 9 answers. Harbour Leadership is not listed on either.
Observation: The answers draw on third-party directories. The client’s own service page was never cited.
Inference (medium confidence): The service page opens with 180 words about coaching philosophy. Location, audience, and services appear only in the lower half. Combined with the missing directory listings, no easily retrievable passage matches the question.
The audit cannot see Perplexity’s retrieval process, so this is a likely cause, not a confirmed one.
Fixes:
- Rewrite the first two sentences of the service page to state who is coached, what the service is, and where.
- Apply for listings on both directories.
- Add ProfessionalService schema with sameAs links to LinkedIn and each directory listing once live.
Verification: Rerun the same nine answers, same wordings, four weeks and eight weeks after the fixes go live. The working threshold for success is a citation in 3 or more of 9 answers in both reruns.
If mentions rise but citations do not, the directory listings may be contributing to visibility without producing citations to the client’s own site. If nothing moves, revisit the diagnosis and check whether the original assumptions still hold.
That verification step carries one honest caveat.
Engines change their sources every week regardless of what you do, so a re-test shows whether results moved after the fix, not proof that the fix caused it. A report that claims otherwise is overselling its own data.
FAQ
How many times should I run a prompt to check whether ChatGPT cites my website?
More than once, and in more than one wording. A reasonable starting point is three wordings of each question, run three times each. Treat the result as directional, not precise.
What is the difference between being mentioned and being cited by AI?
A mention means the AI names your business in its answer. A citation means it links your page as a source. Track them separately, because well-known brands are often mentioned without being cited.
Why does ChatGPT give different sources when I ask the same question twice?
AI answers are generated with built-in randomness, and retrieval results change over time. One analysis found only 2.2% of ChatGPT citations remained after three runs of the same prompt.
Should an AI visibility audit combine ChatGPT, Perplexity and Gemini into one score?
No. Each engine cites different sources, so a combined score hides where you are visible and where you are not. Report each engine separately.
Can schema markup guarantee that ChatGPT or Perplexity will cite my site?
No. Schema helps machines understand who you are and which profiles belong to you, but no AI lab has confirmed it directly controls citation. Anyone guaranteeing citations is guessing.
How long after fixing my site should I re-test AI citations?
Wait at least four weeks, then compare more than one re-test. Weekly source changes are large enough that a single snapshot can mislead in either direction.
A screenshot tells you what an AI said once. It cannot tell you what to fix, or whether the fix worked.
That gap is the audit.
If you want to know what AI systems actually say about your business, AI Visibility Studio runs audits built this way: real buyer questions, repeated runs, and a re-test plan for every fix.
Originally published on Medium ↗