How to Build an AI Citation Benchmark: Prompts, Competitors, and a Weekly Measurement You Can Actually Keep
Most attempts to measure AI visibility die of ambition. People try to track every prompt, every engine, and every competitor, produce a spreadsheet nobody opens twice, and conclude the whole thing is unmeasurable. It is measurable. It just needs a benchmark small enough to run every week and honest enough to survive the questions a sceptic would ask. This is the framework we use, built on Microsoft Clarity’s free reporting, which as of summer 2026 can give eligible projects two of the three measurement layers you need without paying for a separate AI visibility platform.
Key Takeaways
- A benchmark needs three parts: a fixed set of prompts that represent how buyers ask, a stable measurement source, and a reading taken the same way every time.
- Microsoft Clarity’s Citations dashboard, generally available since May 13, 2026, reports observed citation activity: page citations, share of authority, grounding queries, cited pages, AI referral traffic, and trendlines.
- Clarity’s Topic Insights, launched July 9, 2026 and still in beta, adds controlled prompt-based benchmarking inside the same tool: you define topics, prompts, and competitors, and it reports who gets cited. Microsoft says to treat it as directional.
- Share of authority is the competitive metric, and Microsoft documents that its daily calculation can run higher than broader measures, so read movement, not the raw number.
- Clarity is not a complete census of every AI platform, so a third, lighter layer of direct checks in specific engines still earns its place.
Why Do Most AI Visibility Measurements Fail?
Because they measure everything once instead of a few things repeatedly. A one-off audit that checks a hundred prompts across five engines gives you a large number and no way to interpret it, since you have nothing to compare it to and no way to repeat it identically next month.
The other failure is the opposite: relying on a single vanity check. Someone types their brand into ChatGPT, sees themselves mentioned, and calls it visibility. That tells you nothing about the questions buyers ask when they do not already know your name, which is where the citations that matter actually happen.
A benchmark fixes both problems by being small, fixed, and repeated. You choose the prompts once, you measure the same way each time, and the value comes from the line over time rather than any single reading.
The Three Layers, Because They Answer Different Questions
Keep these separate in your head and your spreadsheet, because blending them is how measurement turns into wishful thinking.
The first layer is observed citation data: what Clarity’s Citations dashboard records about how AI systems actually retrieved and cited your pages in supported experiences. You did not define the grounding queries being reported. They came from the retrieval activity Clarity observed, which is what makes this layer different from a benchmark you designed yourself.
The second is controlled benchmarking: Clarity’s Topic Insights, where you define the topics, prompts, and competitors, and the tool runs them and reports who was cited. You chose the questions, so it measures your hypothesis about what matters, consistently, against named competitors.
The third is direct engine checks: running a handful of prompts yourself in specific consumer AI products your buyers use. These are observations, not measurements. They vary by session, they are not aggregated, and they exist to fill the gap the first two layers cannot see.
Step One: Choose a Prompt Set That Represents Buyers, Not Your Brand
Start with the questions someone asks before they know you exist. Our working rule is ten to fifteen prompts for each core topic, which happens to match how Topic Insights is structured, and they should mirror how people actually phrase things to an AI, conversational and specific, not keyword fragments.
Build them from three sources. The questions clients ask in first calls and emails, which are the truest record of buyer language you own. The grounding queries Clarity shows once it has data, since those are the retrieval phrasings AI systems actually used to reach you. And a small set of unbranded comparison prompts, the “how do I choose a provider for X” and “what should X cost” questions where competitors get cited.
Then freeze the set. The temptation to keep tweaking prompts is the same temptation that ruins every benchmark, because a changed question is a new measurement, not a trend. Add a prompt only when a new buyer question clearly earns a place, and mark the week you added it. Keep the set stable across layers too: use the relevant prompts within Topic Insights and the same prompts for your direct checks wherever possible, so the layers remain comparable over time.
Step Two: Read the Observed Citation Metrics in the Right Order
Clarity’s Citations dashboard reports six things, and they answer different questions, so read them as a sequence. Microsoft’s documentation defines them precisely, which is worth taking literally.
Page citations first: the count of times your pages were referenced in AI answers over the period. That is your citation volume. It plays a similar diagnostic role to impressions in traditional search measurement, but it is not an impression count, and Microsoft is explicit that a citation count says nothing about how prominently you appeared in the answer.
Then share of authority: your domain’s percentage of total citations across the citation activity included in Microsoft’s daily calculation. That is your competitive position over the observed data, and it is the number to watch over time. Then the diagnostic pair. My cited pages shows which URLs earn citations. Grounding queries shows the retrieval phrasings behind them, which are often not what a person typed, and Clarity now groups them into topics and flags branded queries separately, so you can split “people already looking for us” from “people looking for what we do.” AI referral traffic, the share of sessions arriving from AI assistants, comes last, because it is a downstream number that often stays small even when citations grow.
Step Three: Handle Share of Authority Honestly
It is the best metric in the dashboard and the easiest to misread. Microsoft calculates it daily, based on the queries where your domain participated, and its own documentation notes this can produce higher values than a broader market-wide calculation would. So a share of authority figure is not your slice of the whole market. It is your slice of the citations on the days and queries where you already showed up.
Two practical rules follow. Compare movement within the same report over time rather than treating any single percentage as an absolute score, and read share alongside raw citation counts, because share can hold flat while volume climbs, which means the whole space is growing and you are keeping pace rather than gaining.
One scoping note. The core Citations view does not hand you a universal competitor leaderboard. For named competitors, you move to the second layer.
Step Four: Run Controlled Benchmarks in Topic Insights
This is the layer that used to require a paid tool or a manual slog, and since July 9, 2026 it is inside Clarity. Topic Insights lets you create topics, define representative prompts for each, and name the competitors you want benchmarked. It then evaluates the AI-generated responses, identifies which domains and pages were cited, measures how much each source contributed, and aggregates the results at the topic level, so you get a consistent view rather than isolated anecdotes.
It also addresses the competitor question directly, since Microsoft confirms it shows which domains are cited alongside you and lets you define competitors. In hands-on testing, Topic Insights also surfaces domains beyond the competitor set you define, which can reveal sources you did not initially think to track.
Read the disclaimer before you read the output, because Microsoft states it plainly. Topic Insights is in beta, currently uses GPT-5.3 grounded by WebIQ, Microsoft’s own retrieval layer, allows ten reports per week per project, and should be used for directional monitoring only. In practice, that makes it a view of what GPT-5.3 grounded through Microsoft’s WebIQ retrieval layer cites, rather than a direct measurement of what ChatGPT or Gemini is showing users. Practitioners also note attribution sits at the topic level rather than the individual prompt, you cannot read the actual generated answer, and reports are not yet linked by a trendline, so comparing runs means opening the older report and reading the numbers side by side.
None of that makes it less useful. It makes it a benchmark, which is exactly what you want: same prompts, same competitors, same method, repeated.
Step Five: Add Direct Checks Where the Tool Cannot See
Clarity does not represent a complete census of every answer generated by every AI platform. Its citation reporting covers supported AI experiences, and Topic Insights runs on one model over one grounding layer, so when you need to know what a specific consumer-facing engine is actually showing your buyers, check it directly.
Keep this layer light or it will not survive. Our working rule is five to ten prompts from the frozen set, in two or three engines your buyers actually use, and a simple table of which domains were named. You are not building a rival dataset. You are checking the specific products the tool cannot see, and putting names to what you find there.
Be honest with yourself about what this layer is. Hand-run prompts vary by session, location, and phrasing, and they are observations rather than measurements in the way Clarity’s aggregated data is. Log them as observations and let the first two layers carry the trend.
Step Six: Take the Reading the Same Way Every Time
Same day, same window, same order, then stop. Open Citations, note page citations and share of authority for the week, list any pages that entered or dropped out of the cited-pages view, and skim the grounding queries for phrasings you have not seen before. Then run or review your Topic Insights report for the week and note citation rate, share of authority, and any competitor movement per topic. Then run the direct checks and log the cited domains. Our rule of thumb is roughly an hour, and it should not grow.
The output is three lines that only mean something over time. Is share of authority rising, flat, or falling on the questions that matter? Which of your pages are earning and losing citations, and which topics are competitors winning? And what do the direct checks show in the engines that matter most to you?
Our rule is eight weeks before you judge anything. After that you have a real benchmark: not a score, but a direction, tied to specific pages, topics, and questions, that tells you where the next piece of work should go. That is worth more than any audit, because you can act on it and then watch whether the action worked.
Frequently Asked Questions
Q: How do I measure whether AI systems cite my website?
A: Microsoft Clarity’s free Citations dashboard, generally available since May 2026, reports page citations, share of authority, grounding queries, cited pages, and AI referral traffic for supported AI experiences. Its Topic Insights feature, in beta since July 2026, adds controlled prompt-based benchmarking against named competitors. Pair both with a few direct checks in the engines your buyers use.
Q: What is share of authority in Microsoft Clarity?
A: It is the percentage of citations attributed to your domain compared with other cited domains across the same queries where your domain appeared. Microsoft notes its daily calculation can run higher than broader measures, so use it to track movement over time rather than as an absolute market share.
Q: Does Microsoft Clarity show which competitors AI is citing?
A: The core Citations view does not give you a universal competitor leaderboard, but Clarity’s Topic Insights feature identifies domains cited alongside you within the topics you define and lets you name competitors for direct comparison. In hands-on testing it has also surfaced domains beyond the defined competitor set, which can reveal sources you had not thought to track.
Q: What is Topic Insights in Microsoft Clarity and how reliable is it?
A: Topic Insights lets you define topics, prompts, and competitors, then reports which domains and pages AI answers cited and how much each contributed. Microsoft says it is in beta, runs on GPT-5.3 grounded by WebIQ, allows ten reports per week per project, and should be used for directional monitoring rather than as a guarantee.
Q: How many prompts should an AI citation benchmark include?
A: Our working rule is ten to fifteen for each core topic, which matches how Topic Insights is structured, drawn from real client questions, Clarity’s grounding queries, and unbranded comparison questions. Freeze the set so changes over time reflect your visibility rather than changes to the questions.
Q: How often should I check my AI citation data?
A: Weekly, at the same time and in the same order, is our recommendation. The value of a benchmark comes from the trend across weeks, and Microsoft itself recommends establishing a baseline and monitoring changes, so a short, repeatable reading beats a long occasional audit.
Almost nobody measures AI visibility because almost everybody tries to measure all of it. Measure a little of it, the same way, every week, across the three layers, and within two months you will know something about your visibility that a hundred-prompt audit could never tell you: which way it is moving, and why.
AI Visibility Studio helps websites structure content so AI systems can find it, understand it, cite it, and actually use it when generating answers. aivisibilitystudio.com
References
- Microsoft Learn, “Citation Dashboard Overview” (Clarity AI Visibility): https://learn.microsoft.com/en-us/clarity/ai-visibility/ai-citations
- Microsoft Clarity Blog, “Citations in Microsoft Clarity Now Generally Available” (May 13, 2026): https://clarity.microsoft.com/blog/citations-now-generally-available/
- Microsoft Clarity Blog, “Introducing Topic Insights: Actionable Content Recommendations” (July 9, 2026): https://clarity.microsoft.com/blog/topic-insights-announcement/
- Microsoft Advertising Blog, “Understand how humans and AI engage with your brand with Microsoft Clarity” (July 2026): https://preview-about.ads.microsoft.com/en/blog/post/july-2026/understand-how-humans-and-ai-engage-with-your-brand-with-microsoft-clarity
- Aubrey Yung, “Microsoft Clarity Topic Insights: A Hands-On Review” (July 9, 2026; practitioner limits on attribution and trendlines): https://aubreyyung.com/observatory/microsoft-clarity-topic-insights-first-impressions/
Originally published on Medium ↗