How to measure AI visibility: prompts, share of voice, citation rate
Measure AI visibility by hand: 10 buyer questions, 7 engines, 4 metrics (mention rate, citation rate, share of voice, accuracy), and how to read the noise.
LLaunchScaler·Published ·8 min read
Measure AI visibility by running a fixed set of about 10 buyer questions through ChatGPT, Perplexity, Claude, Google AI Overviews, AI Mode, Gemini and Copilot, then recording four things per answer: whether you are mentioned, whether your site is cited with a link, which competitors are named, and whether the description of you is accurate. Those give you a mention rate, a citation rate, share of voice and an accuracy rate.
There is no standard "AI visibility score." Every tool that sells one computes it its own way from numbers like these. Computing them yourself takes an afternoon a month and tells you what the score would hide. Below is the method, a worked example, and how to read the noise between runs, checked in September 2026.
What should you measure in AI answers?
Measure four rates across a fixed set of questions and engines: how often you are named, how often you are linked, what share of all brand mentions are yours, and how often what is said about you is right. Position within an answer is worth recording, but it is too unstable to track on its own.
Metric
Formula
What it tells you
Mention rate
Answers that name you ÷ answers collected
Whether engines think of you for these questions at all
Citation rate
Answers that link your domain ÷ answers collected
Whether engines use your pages as a source
Share of voice
Your mentions ÷ mentions of you plus named competitors
Questions, answered
What people ask about this
01
What is an AI visibility score?
There is no standard one. Each tool defines its own, usually from how often a brand is mentioned or cited across a set of prompts. You can compute the underlying numbers yourself: mention rate, citation rate, share of voice against competitors, and accuracy.
Answers that describe you with no material error ÷ answers that name you
Whether the mentions help or hurt
Position (secondary)
Your rank in any list the answer gives
A rough hint only; it changes run to run
Mention and citation are separate on purpose. An answer can name you without linking you, because it learned about you from other sites, or link one of your pages without recommending you. Both matter, and they call for different fixes.
How do you choose the questions?
Write the questions a buyer asks before they know your name. Those are the questions where a recommendation is won or lost, and they are the ones your brand-name searches will never show you. Aim for about 10, covering different intents, and keep them fixed from month to month so the numbers are comparable.
Cover these intents, one or two questions each:
Best-of: "What is the best invoicing software for freelance designers?"
Use case: "How can I automate payment reminders for my clients?"
Alternatives: "What are alternatives to [the market leader] for small studios?"
Comparison: "[Competitor A] vs [Competitor B] for freelancers?"
Constraint: "Cheapest invoicing tool that supports multiple currencies?"
Problem: "Clients keep paying late, what tools help?"
Leave your brand name out of every question. A branded question ("is Beacon Invoicing good?") measures something else: accuracy, not visibility. Run those separately when you check how engines describe you.
How do you run the questions in each engine?
Run each question in each of the seven engines, the same way every month, and paste every answer into a sheet. Control the settings that change answers: search mode, sign-in state, memory and location. OpenAI says ChatGPT uses an approximate location from your IP address and, with memory on, saved memories when it rewrites a search query.
Engine
How to run it
Settings to hold constant
ChatGPT
Turn on Search from the tools menu, then ask
Same account state; memory off or a temporary chat
Perplexity
Ask in the default search
Same mode each month
Claude
Ask with web search enabled
Same model choice
Google AI Overviews
Search on google.com; record whether an AI Overview appears
Signed out, same country
Google AI Mode
Ask in AI Mode
Signed out, same country
Gemini
Ask in the Gemini app
Same account state
Microsoft Copilot
Ask in Copilot
Same account state
AI Overviews do not appear for every search, so record "no AI Overview" as its own outcome rather than as a miss. Google also tells users its AI Overviews "can and will make mistakes," which is one more reason to record accuracy, not only presence.
Use one row per answer, with these columns: date, engine, question, mentioned (yes or no), linked (yes or no), position if listed, competitors named, cited URLs, accuracy note. Ten questions across seven engines is 70 rows a month.
How do you calculate share of voice?
Count every brand mention across all 70 answers, yours and the competitors', then divide your count by the total. Only include competitors you actually lose deals to, named in advance, so the denominator does not change with whatever random brands an engine lists that month.
A worked month for a hypothetical invoicing tool:
Answers collected: 70 (10 questions × 7 engines).
Answers that name you: 14. Mention rate: 14 ÷ 70 = 20%.
Answers that link your domain: 6. Citation rate: 6 ÷ 70 = 8.6%.
Mentions of Competitor A: 30. Competitor B: 21.
Share of voice: 14 ÷ (14 + 30 + 21) = 21.5%.
Answers naming you with no material error: 11 of 14. Accuracy rate: 79%.
Split each rate by engine as well. Engines draw from different sources: a Loganix compilation reports that only 11% of cited domains were shared by ChatGPT and Perplexity in one large analysis. A 40% mention rate in Perplexity and 0% in ChatGPT is a different problem from 20% in both.
How much do AI answers vary between runs?
A great deal, for lists. SparkToro had 600 volunteers run 12 prompts through ChatGPT, Claude and Google's AI 2,961 times, and found "there's a <1 in 100 chance that ChatGPT or Google's AI, if asked 100X, will give you the same list of brands in any two responses," and roughly a 1 in 1,000 chance of the same order.
The same study found what is stable: how often a brand appears across many runs. One hospital appeared in 69 of 71 ChatGPT answers to the same question, a 97% visibility rate, while being the first recommendation in only 25 of them. SparkToro's conclusion was that ranking positions are close to meaningless, but "Measuring that percent visibility is (probably) a reasonable way" to gauge presence.
That shapes how you read your sheet:
Trust rates over many answers, not any single answer. One run of one question tells you almost nothing.
Treat position as a hint. Report it, but do not set goals on it.
Compare monthly totals, and act on changes that hold for more than one month.
If a question matters a lot, run it several times in the same month and count appearances. SparkToro suggests that knowing an AI's recommendation set really takes "at least 60-100X."
Microsoft makes the same point about Copilot. Its AI Performance help says citation counts change "based on user demand, content freshness, model updates, partner refresh cycles," and that its Compare view "shows what changed between two periods" but "does not explain why."
What do Google, Bing and GA4 report directly?
Each covers one slice. Search Console's generative AI performance report counts impressions of your links in AI Overviews and AI Mode, rolled out to all sites on August 31, 2026. Bing Webmaster Tools' AI Performance report counts Copilot citations and grounding queries. GA4 labels clicks from assistants such as ChatGPT in its AI Assistant channel.
Source
What it measures
What it misses
Search Console generative AI performance report
Impressions of links to your site in AI Overviews and AI Mode, by page, country and device
ChatGPT, Perplexity, Claude, Copilot; mentions without a link
Bing Webmaster Tools AI Performance
Total citations, average cited pages, grounding queries, and Citation Share per grounding query
Every engine outside Microsoft
GA4 AI Assistant channel (since May 13, 2026)
Sessions that arrive from recognized assistants
Answers that never produce a click; visits that arrive with no referrer
Your own prompt runs
Mentions, citations, competitors and accuracy in all seven engines
Scale; it is a sample
Bing's Citation Share is the closest first-party equivalent to share of voice: "the percentage of citations attributed to your site out of all citations shown for a specific grounding query." It does not name the other sites. Setting up the GA4 side is covered in how to track ChatGPT and Perplexity traffic in GA4.
When is it worth paying for an AI visibility tool?
When you need more questions, more runs or daily tracking than a monthly sheet can hold. Tools automate the runs and compute the rates, and some add prompt research. They differ widely in engines covered, prompts per month and price; a comparison of what each one checks is in best GEO and AEO tools, and trackers focused on ChatGPT are compared in ChatGPT visibility tracker tools.
Before paying, run the manual method for one month. It shows you whether you have a visibility problem, an accuracy problem or an access problem, and a tool only helps with the first. Ongoing tracking after that is covered in AI visibility monitoring.
What should you do with the results?
Work from the answers you lose. For each question where a competitor is named and you are not, open the sources the engine cited and ask whether you can get onto them. For each inaccurate mention, find the page the error came from. For low citation rates with decent mention rates, look at whether engines can reach and read your pages.
LaunchScaler's full audit runs this measurement on your domain: 10 buyer questions put to 7 answer engines (ChatGPT, Perplexity, Claude, Google AI Overviews, Google AI Mode, Gemini and Copilot), all 70 answers read. It reports buyer-question coverage, which competitors each answer names, your citation share against them and whether engines describe your brand accurately, alongside twenty backlink checks. It costs $19 once for your domain and unlocks on your report after the free scan. Run the free scan first, then open the full audit.
Write the questions a buyer asks before they know your name, run them in ChatGPT with search on, and record for each answer whether you are named, whether your site is linked, which competitors are named and whether the description of you is correct. Repeat monthly with the same questions.
03
Why do AI answers change every time I ask?
Because generated answers are not fixed rankings. SparkToro's study of 2,961 runs found less than a 1 in 100 chance that ChatGPT or Google's AI gives the same list of brands twice in 100 runs. How often a brand appears across many runs is far more stable than its position in any one answer.
04
Which AI engines should I measure?
Cover the seven major answer surfaces: ChatGPT, Perplexity, Claude, Google AI Overviews, Google AI Mode, Gemini and Microsoft Copilot. Engines cite different sources, so a result in one does not predict the others.
05
Do Google and Bing report AI visibility?
Partly. Search Console's generative AI performance report counts impressions of your links in AI Overviews and AI Mode, and Bing Webmaster Tools' AI Performance report counts citations in Copilot and shows your citation share for grounding queries. Neither covers ChatGPT, Perplexity or Claude.
Search Console now lets you pull a site out of AI Overviews and AI Mode together. Page-level options, what each costs, and why Google-Extended isn't one.