What is AI visibility, and how do you actually measure it?
Updated August 2026 6 min readAI visibility is how often, and how favourably, an AI engine names your brand when someone asks a question in your category. It is not a ranking and it does not follow from one. Across 15,000 long-tail queries, only about 12% of the URLs cited by AI assistants also ranked in Google's top 10 for the same prompt[2]. Your Google position tells you almost nothing about whether ChatGPT recommends you.
Why the question is suddenly worth money
Because the answer increasingly is the destination. In the first four months of 2026, 68.01% of US Google searches ended without a single click to any website, up from 60.45% in 2024[1]. The research still happens. It just resolves inside a summary, and the brands named in that summary get the consideration.
The visits that do arrive from AI assistants behave differently too. Semrush puts the average AI search visitor at 4.4 times the value of a traditional organic visitor[3], and Adobe's 2025 retail data showed AI referrals bouncing 27% less often and staying 38% longer than non-AI traffic[4]. Fewer visitors, further down the decision.
What AI visibility is not
Three confusions cost people the most time:
- It is not a rank. There is no position 1 inside a generated answer. A brand is either named or not, and if it is named it may or may not carry a link. Tools that report a "rank" are ranking you against other brands in their own sample, not reporting something the engine holds.
- It is not your Google ranking in a new outfit. The overlap is small and varies sharply by engine (table below).
- It is not traffic. Most AI visibility never becomes a click. That is the point of it, and it is why measuring only sessions understates the channel badly.
How much does your Google ranking actually carry over?
Not much, unless you are talking about Google's own AI Overviews. Ahrefs ran 15,000 long-tail queries in July 2025 and checked how often an AI-cited URL also sat in Google's top 10 for the same prompt[2]:
| Engine | Overlap with Google top 10 |
|---|---|
| ChatGPT (in-text citations) | 8.0% |
| ChatGPT (reference list) | 6.1% |
| Copilot | 8.2% |
| Gemini | 8.6% |
| Perplexity | 28.6% |
| Google AI Overviews | 76% |
| Average across assistants | 11.9% |
Read the two ends of that table carefully, because they imply different work. AI Overviews summarise what Google already ranks, so classic SEO carries over there. The standalone assistants mostly cite pages that are not in the top 10 at all, often on domains you do not own: a forum thread, a review roundup, a comparison article someone else wrote.
The four metrics worth tracking
Anything beyond these four is usually a vendor inventing a proprietary score so its dashboard cannot be compared with anyone else's.
| Metric | The question it answers | How it is computed |
|---|---|---|
| Mention rate | How often are we named at all? | Answers naming your brand, divided by total answers for your prompt set. |
| Share of voice | Named how often relative to competitors? | Your mentions divided by all brand mentions across the same answers. |
| Citation rate | Are our pages the evidence, or is someone else's? | Average number of URLs from your domain cited per answer. |
| Sentiment | When we are named, how are we described? | Scored per mention. Being named as the expensive option is not a win. |
Mention rate and share of voice move independently, and the gap between them is the useful signal. Rising mentions with flat share of voice means the whole category is getting more answer space and you are keeping pace. Flat mentions with rising share of voice means competitors are dropping out. Only one of those is worth a press release.
Citation rate is the metric most people skip and the one that explains everything else. If engines describe you accurately but cite a competitor's comparison page as the evidence, the fix is not more content on your site. It is getting into the sources the engine already trusts.
Why one ChatGPT check tells you nothing
Ask the same question twice and you can get two different answers, with different brands in them. Generative engines sample, so any single run is one draw from a distribution, not a reading.
The arithmetic is unforgiving. Suppose your true mention rate for a prompt is 30%. Run that prompt 30 times and the 95% confidence interval around your measurement is roughly 14% to 46%. Run it 300 times and it tightens to about 25% to 35%. Anything less than a few hundred observations per tracked prompt cannot detect the size of change you would actually act on, which is why a founder checking ChatGPT once a week and feeling encouraged or panicked is reading noise both times.
Read more on sampling
Three practical consequences. First, run prompts on a schedule, daily if you can, and compare rolling windows rather than single days. Second, hold the prompt set still. Adding or removing prompts changes what you are measuring, so a visibility jump that follows a prompt-set edit is an artefact, not a result. Third, log the full answer text, not just a yes or no on the mention. The wording is where you find the reason.
Engines also personalise and localise. Measurement should be run from a consistent, logged-out context, and if your market is regional you need a second run from that region rather than an assumption.
How do you build a prompt set that reflects real buyers?
The prompt set is the measurement instrument. Get it wrong and every number downstream is decorative.
- Write prompts your buyers would type, not your keywords. People ask assistants full questions with constraints attached: "best CRM for a 12-person agency that needs Xero integration", not "best CRM".
- Cover the whole funnel. Category discovery, shortlist comparison, objection handling ("is X worth it", "X alternatives"), and one or two branded prompts to catch reputation problems.
- Keep branded prompts out of the headline number. An engine will name you when the question already contains your name. Counting those inflates visibility and hides the problem.
- Fix the set for at least a quarter. Comparability over completeness.
- Aim for 30 to 60 prompts. Below 30 the category is not covered. Above 60 the cost rises faster than the insight.
How do you see AI traffic in your own analytics?
Partially, and knowing which part is missing matters. Assistant referrals arrive with identifiable referrers: chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com, claude.ai. Build one GA4 segment or channel group from that list and you can see AI-sourced sessions, their conversion rate and their landing pages.
What that segment cannot show you: traffic from Google AI Overviews and AI Mode, which arrives labelled as ordinary Google organic, and the far larger share of AI visibility that produces no click at all. So treat the GA4 number as a floor, not a measurement. It is the only part of AI visibility your analytics can see, and it is the smallest part.
What does a good number look like?
There is no universal benchmark, and anyone quoting one is selling something. Share of voice is defined against your competitive set, so 15% might be dominant in a fragmented category and invisible in one with three players.
Three honest reference points instead. In most categories the leading brand takes a minority of total mentions rather than a majority, so aiming to "own" the category is the wrong target. A mention rate that moves more than a few points week to week is usually a measurement artefact rather than a real change. And first movement from serious work tends to show up in 6 to 10 weeks, because it depends on third-party sources being published, indexed and then retrieved.
The number that matters is your own baseline, measured properly, then tracked against a competitor set you actually lose deals to. If you want that baseline without building the pipeline yourself, our free AI visibility audit runs your prompt set across five engines and shows you which pages the answers are built from. If you would rather do it in-house, everything above is the method, and what GEO is covers what you do with the findings.
Frequently asked
What is AI visibility?
AI visibility is how often, and how favourably, AI engines such as ChatGPT, Perplexity, Gemini and Google AI Overviews name your brand when people ask questions in your category. It is measured across a fixed set of buyer prompts, run repeatedly, rather than by checking a single answer once.
How is AI visibility different from SEO rankings?
A ranking is a position in a list of links; AI visibility is presence inside the generated answer. They overlap far less than people assume. Across 15,000 queries, only about 12% of AI-cited URLs also ranked in Google's top 10 for the same prompt, though Google's own AI Overviews are the exception at roughly 76%.
How do you measure AI visibility?
Fix a set of 30 to 60 buyer prompts, run them across the engines your market uses on a daily schedule, and record four things: mention rate, share of voice against a named competitor set, citation rate for your own URLs, and sentiment. Compare rolling windows, not single days, because engines are non-deterministic.
How many times do I need to run a prompt for the number to mean anything?
Enough that the confidence interval is smaller than the change you care about. At a true 30% mention rate, 30 runs give you roughly 14% to 46% at 95% confidence, while 300 runs give roughly 25% to 35%. One run tells you nothing you should act on.
Can you pay an AI engine to recommend your brand?
Not for the organic part of the answer. Engines are adding advertising surfaces around answers, but the brands named inside a recommendation are selected from retrieved sources. That is why the work is about what those sources say, not about buying placement.
Sources
- Search Engine Land, "Google zero-click searches reach 68% in early 2026", 9 June 2026, reporting SparkToro's analysis of Similarweb US clickstream data for January to April 2026.
- Ahrefs, "Only 12% of AI Cited URLs Rank in Google's Top 10 for the Original Prompt", study of 15,000 long-tail queries, July 2025.
- Semrush, "AI SEO statistics", citing Semrush's 2025 analysis putting the average AI search visitor at 4.4 times the value of a traditional organic visitor.
- Adobe, 2025 retail analytics on AI-referred visits (27% lower bounce rate, 38% longer visits), as reported by Semrush.
Find out if AI recommends you.
The free audit shows which engines mention you, which recommend competitors instead, and the exact pages behind those answers.
Get the free audit48 hours · no call · no mailing list