AI visibility tools compared: what they sample, what they cost, and what they cannot tell you
Updated September 2026 7 min readAI visibility tools do not measure the same thing, so their numbers will never match: each one samples a different set of prompts, engines, runs and collection methods. In a test of 1,000 prompts published in September 2026, the brands an engine named through its API overlapped with the brands shown in the real interface only 15.5% to 23.8% of the time[2]. Pick a tool by what it samples, not by its dashboard.
What does an AI visibility tool actually measure?
A sample of answers, never the whole picture. Every tool in this category does the same basic job: it sends a list of prompts to a set of AI engines on a schedule, stores the answers, and counts which brands and which source pages appear. The differences that matter are in four settings: which prompts, which engines, how many runs, and whether the answer comes from the engine's API or from the interface a real buyer sees.
Those settings matter because the engines are not stable. SparkToro and Gumshoe had 600 volunteers run 12 prompts through ChatGPT, Claude and Google's AI a combined 2,961 times in November and December 2025. The chance of getting the same list of brands twice was under 1 in 100, and the same list in the same order was closer to 1 in 1,000[1]. Their conclusion is the one we work from: visibility measured as a percentage across many prompts and many runs is a reasonable metric, and a single "position" is not[1].
If the metrics themselves are new to you, start with our explainer on what AI visibility is and how to measure it, then come back to the tools.
What does each tool sample, and what does it cost?
Here is what each vendor's own pages said in September 2026. Prices are as listed by the vendor, so most are in US dollars. We use Peec and Ahrefs Brand Radar ourselves; we have not held paid accounts on the other three.
| Tool | Prompts you control | Engines | Run cadence | Entry price |
|---|---|---|---|---|
| Peec AI | 50 (Starter), 150 (Pro), 350 (Advanced) | Choose 3 models on self-serve plans; up to 13 on Enterprise | Daily | Not in the page's static HTML, so we have not quoted it[4] |
| Ahrefs Brand Radar | 5, 10 or 20 custom prompts bundled with Ahrefs Lite, Standard, Advanced; plus a 448M-prompt index | AI Overviews, AI Mode, ChatGPT, Perplexity, Gemini, Copilot; Claude on custom prompts | Custom prompts daily, weekly or monthly; index re-tested monthly | Custom prompts from $50/mo; AI Visibility Index from $199/mo[6] |
| Semrush AI Visibility Toolkit | 25 custom prompts on the base plan | ChatGPT, Google AI, Gemini, Perplexity | Prompt tracking daily; Brand Performance weekly | €94.94/mo per domain, billed annually[7] |
| Otterly.AI | 15 (Lite), 100 (Standard), 400 (Premium) | ChatGPT, AI Overviews, Perplexity, Copilot; Claude, Gemini, AI Mode as paid add-ons | Daily | $29/mo Lite, $189/mo Standard, $489/mo Premium[5] |
| Profound | Custom; free trial of 50 prompts for 7 days | Up to 9 on Enterprise, including ChatGPT, Perplexity, AI Mode, Gemini, Copilot, Claude, AI Overviews | Daily | Enterprise price not published[9] |
Three things in that table decide more than the price does.
- Prompt count per euro. Otterly's Lite tier is the cheapest way in, but 15 prompts do not cover a category. Semrush's base plan gives you 25. Peec starts at 50.
- Engine caps. Peec's self-serve plans let you choose 3 models at every tier[4]. Otterly charges extra for Claude, Gemini and AI Mode[5]. If your buyers live in one engine you cannot track, the dashboard is measuring someone else's market.
- Whose prompts. Ahrefs' index is huge, but every prompt in it is "modeled from real user searches in Ahrefs' keyword database"[6]. That is keyword data rewritten as questions, not a log of what people typed into ChatGPT.
What can't each tool tell you?
Every one of them has a blind spot, including the two we pay for.
- Peec AI gives you daily runs on prompts you write, which is the right shape for tracking. The limits: 3 engines on self-serve plans, and a "position" metric that sits next to visibility in the overview. Given the SparkToro finding that list order almost never repeats[1], we read visibility and share of voice from Peec and treat position as colour, not evidence.
- Ahrefs Brand Radar is the only tool here that shows you a market you have not set up prompts for. The limit is freshness and origin: the index's question sets are "re-tested monthly on a 90-day reporting window"[6], and the prompts are modelled from search keywords. It is good for "who owns this category" and weak for "did last week's placement move anything".
- Semrush reports an AI Visibility Score defined as your mentions "compared to the median number of mentions for your top industry competitors"[8]. That makes the score relative to a competitor set Semrush picks, so it cannot be compared with any other tool's number.
- Otterly is the cheapest honest entry point, but its base engine list leaves out Gemini and AI Mode unless you pay for add-ons[5].
- Profound covers the most engines on paper[9], but with no public price you cannot judge value until a sales call, which rules it out for most small teams.
None of the five pricing pages we checked says plainly whether answers are collected from the engine's API or from its consumer interface. Ask before you buy, for the reason in the next section.
Why do two tools give different numbers for the same brand?
Because they are measuring different samples, and at least four variables differ between them.
- API versus interface. Surfer ran 1,000 prompts across five AI products on 4 August 2026, once through APIs and once by scraping the interfaces, collecting 13,779 answers. Brand overlap between the two methods was 15.5% to 23.8%, rising to 21.3% to 31.6% after merging spelling variants. Through the API, ChatGPT named 13.8 brands per answer against 7.9 in the interface, and cited 3.1 sources against 12.1[2]. Two caveats: Surfer sells a scraping-based tracker, and it took one sample per prompt, so it says itself that some of the gap is ordinary randomness rather than the channel.
- Prompt wording. Peec's study of 37,804 responses across 1,754 prompts found ranking-style prompts added about 20% average visibility over open-ended questions, and short keyword-style prompts up to 25% more than persona-style ones[3]. Two tools with different prompt styles will disagree even on the same engine and the same day.
- Run count. SparkToro suggests running a prompt 60 to 100 times to understand an engine's recommendation set[1]. A tool running each prompt once a day reaches that after two or three months. A monthly index takes longer still.
- Metric definitions. Visibility, share of voice and proprietary scores are computed differently by each vendor, as the Semrush definition above shows.
Read more: what to do when your tools disagree
Do not average them. Pick one tool as the system of record for trend lines, keep its prompt set fixed for at least a quarter, and use the second tool only to answer a different question, such as category-wide share from an index. When a number moves, check the raw answers before reporting it. A jump caused by one engine update or one reworded prompt is not a result.
When does a spreadsheet beat a tool?
When the question is small, one-off or qualitative. A spreadsheet is the better choice in three situations.
- You are deciding whether the channel matters at all. Ten buyer prompts, three engines, five runs each is 150 answers. That is an afternoon of copy and paste, and it tells you whether you are absent, occasional or present. You do not need a subscription to learn that you are absent.
- You need to read the answers, not count them. Tools count mentions. A human reading the answer notices that you are named as the expensive option, or that the engine is quoting a two-year-old review. That judgement is where the work starts.
- Your prompt set is tiny. If you care about 10 prompts, a 15-prompt plan buys you a chart, not insight.
A tool wins as soon as you need trend lines. Peec calculates monthly answers as prompts times models times tracking frequency[4], so its smallest plan running daily produces 50 × 3 × 30 = 4,500 answers a month. Nobody collects that by hand, and without that volume the noise SparkToro documented swamps any real change.
Which AI visibility tool should you choose?
The one whose sample matches your buyers. Our rough rule, in September 2026:
- Testing the water on a small budget: a spreadsheet first, then Otterly Lite or Ahrefs' bundled custom prompts if you already pay for Ahrefs.
- Tracking your own prompt set week to week: a daily prompt tracker such as Peec, Otterly Standard or Semrush, with 30 to 60 prompts and the engines your buyers use.
- Mapping a whole category, including brands you have not thought of: an index such as Ahrefs Brand Radar, accepting that it moves monthly.
- Enterprise, many markets, many engines: get quotes from Profound and Peec Enterprise and compare engine coverage line by line.
Whatever you choose, the tool tells you where you stand. It does not change it. That comes from the pages the engines cite, most of which sit on sites you do not own, which is why GEO and SEO need different work. If you would rather see a baseline before committing to any subscription, our free AI visibility audit runs your buyer prompts across the major engines and shows which sources the answers are built from.
Frequently asked
What is the best AI visibility tool?
There is no single best one, because each samples differently. Daily prompt trackers such as Peec, Otterly and Semrush suit week-to-week tracking of prompts you write. Ahrefs Brand Radar's index suits category mapping but re-tests monthly. Choose on engines covered, prompts per plan and collection method, and check that the engines your buyers use are included.
How much do AI visibility tools cost?
In September 2026, entry plans ran from $29 a month (Otterly Lite, 15 prompts) through €94.94 a month billed annually (Semrush, 25 prompts) to $199 a month for Ahrefs' AI Visibility Index. Profound does not publish enterprise prices. Compare cost per tracked prompt and per engine, not headline price, because the limits vary widely.
Why does my AI visibility score differ between tools?
Tools use different prompts, engines, run counts and metric definitions, and some collect answers through APIs while others scrape the interface. Surfer found brand overlap between API and interface answers of only 15.5% to 23.8% in August 2026. Pick one tool as your system of record and track its trend rather than comparing absolute numbers.
Can I track AI visibility without a paid tool?
Yes, for a first look. Run 10 buyer prompts through 3 engines 5 times each and log who is named in a spreadsheet. That is 150 answers and shows whether you appear at all. For trends you need far more runs, because answers change almost every time, so a paid tracker becomes worth it once you are acting on the numbers.
Sources
- SparkToro, "NEW Research: AIs are highly inconsistent when recommending brands or products", Rand Fishkin with Patrick O'Donnell (Gumshoe), 28 January 2026. 600 volunteers, 12 prompts, 2,961 runs across ChatGPT, Claude and Google AI, November to December 2025.
- Surfer, "LLM scraped AI answers vs API results", Paulina Kaleta, 25 September 2026. 1,000 prompts, 5 AI products, 13,779 answers, data collected 4 August 2026.
- Search Engine Journal, study by Malte Landwehr of Peec AI on prompt style and brand visibility, 15 June 2026. 37,804 responses, 1,754 prompts, 5 engines.
- Peec AI, "Pricing", plan limits checked September 2026.
- Otterly.AI, "Pricing", checked September 2026.
- Ahrefs, "Brand Radar", product and pricing page, checked September 2026.
- Semrush, "AI Visibility Toolkit Pricing", checked September 2026.
- Semrush Knowledge Base, "AI Visibility Toolkit", checked September 2026.
- Profound, "Pricing", checked September 2026.
Find out if AI recommends you.
The free audit shows which engines mention you, which recommend competitors instead, and the exact pages behind those answers.
Get the free audit48 hours · no call · no mailing list