GEO measurement

How Accurate Are AI Visibility Trackers? What the Numbers Really Mean

Tools like Profound and Peec AI put a confident number on your screen. Here is how much of that number you can actually trust, and how to read it without fooling yourself.

Quick answer

AI visibility trackers are directionally useful, not precise gospel. Because language models return different answers to the same question run to run, any single "you appear 42% of the time" figure is noisy. What you can trust is the direction of your mention rate over many samples, per engine, over weeks. Treat absolute numbers with skepticism.

Every AI visibility tool sells the same promise: a dashboard that tells you how often ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews mention your brand. The dashboards look precise. A score, a percentage, a rank. The uncomfortable truth is that the underlying system is probabilistic, and precision on top of noise is still noise.

This is not a reason to throw the tools out. It is a reason to read them correctly. Below is what the research actually shows about how repeatable AI answers are, why two credible tools can disagree, and which parts of a tracker report you can take to a client meeting.

<1 in 100

chance two identical runs return the same brand list, across 2,961 prompts on ChatGPT, Claude and Google AI (SparkToro, 2026)

80

distinct outputs from 1,000 runs of one prompt at temperature 0 (Thinking Machines Lab, 2025)

24%

brand overlap between ChatGPT's API and its consumer app over 1,000 prompts (Surfer, via Gumshoe, 2026)

Why the numbers move even when nothing changed

AI answers are nondeterministic by design and, it turns out, by accident too. Run the same prompt twice and you can get a different set of brands, a different order, and different sources. This is the root cause of every accuracy question about trackers.

The scale of the variance is larger than most marketers assume. SparkToro's Rand Fishkin and Gumshoe's Patrick O'Donnell ran 2,961 prompts across ChatGPT, Claude and Google's AI Overviews, repeating each prompt 60 to 100 times. They found less than a 1 in 100 chance that any two runs return the same brand recommendations, and roughly 1 in 1,000 odds of the same list in the same order (SparkToro, January 2026). Fishkin's conclusion was blunt: any tool that reports a "ranking position in AI" is selling a metric the data cannot support (Search Engine Journal, 2026).

You might expect this to disappear at temperature 0, the setting meant to make a model deterministic. It does not. Thinking Machines Lab sampled 1,000 completions of a single prompt at temperature 0 and got 80 unique outputs, with the very first divergence appearing at token 103 (Thinking Machines Lab, September 2025). The cause is not the prompt. It is batch-size dependence in the inference kernels: the same request gets grouped differently depending on server load, and the math shifts slightly each time. For a tracker sampling a live production endpoint, this variance is baked in.

The takeaway

A single sample is a coin flip dressed up as a measurement. This is why serious trackers sample the same prompt many times and report a rate, not a one-shot result.

The API-vs-app problem, the caveat nobody puts on the label

The biggest accuracy gap is not run-to-run noise. It is that most trackers query models through their API, while your buyers use the consumer app, and the two return meaningfully different answers. Any tool measuring the API is a proxy for what customers actually see, not a mirror of it.

Surfer compared 1,000 ChatGPT prompts sent through the API against the same prompts scraped from the ChatGPT web app. The brand overlap between the two was 24%, and the API returned citations only about 75% of the time while the app always did (Surfer, reported by Gumshoe, 2026). The consumer app wraps the model in system instructions, retrieval, and formatting rules that the raw API never sees.

Vendors handle this differently, and it is worth knowing which method your tool uses before you trust its numbers. Broadly, some platforms sample the API for depth and volume, while others scrape or simulate the app interface for fidelity to what users see; neither approach is "correct," they answer different questions (Brainlabs, 2026). Check Profound and Peec AI directly for their current methodology rather than trusting a comparison blog, including ours. If a tracker will not tell you whether its data comes from an API or the app, that silence is the answer.

The engines themselves also disagree with each other, which is why a per-engine breakdown matters more than a blended score. An analysis of 680 million AI citations found only about 11% domain overlap between the sources ChatGPT and Perplexity cite (Averi, 2026). Even within one model, Semrush found only 25.6% of cited domains overlapped between ChatGPT's fast and reasoning modes (Semrush, June 2026). Growth advisor Kevin Indig summed it up in that study: "The brand that wins under minimal reasoning is not the brand that wins under high reasoning. These are two different systems."

AI visibility tracker accuracy: what the numbers measure

AI visibility tracker accuracy comes down to one distinction: a mention rate over many samples is a real, defensible statistic, while a single "did it mention me" check is noise. Trackers earn their value by sampling enough to turn the coin flips into a stable percentage you can trend, not by pretending any single answer is the truth.

Think of it like polling. No pollster claims one phone call predicts an election. They call thousands of people and report a range with a margin of error. A good tracker does the same thing with prompts: it runs each buyer question many times per engine and reports the share of runs that mentioned you. The more samples, the tighter the estimate. This is why sample count is the single most important number a vendor can disclose, and the one they most often hide.

What this means in practice: a mention rate that moves from 20% to 35% over a month of consistent sampling is a real signal worth acting on. A jump from 40% to 44% inside a single week, or a metric built on one check per engine, is likely inside the noise band. The number is only as accurate as the sampling behind it.

What trackers can and cannot tell you reliably

The honest way to use a tracker is to trust it for relative movement and treat its absolute figures as estimates. The table below splits the typical dashboard into signals you can defend and signals you cannot, based on the variance research above.

AI visibility tracker signals, ranked by how much you can trust them
Signal on the dashboardReliable?Why
Your mention-rate trend over weeksYesMany samples average out the run-to-run noise; direction is stable.
Share-of-voice movement vs named competitorsYesA relative comparison inside one sampling method is apples to apples.
Which engines mention you at allYesPresence or absence is far more stable across samples than exact rate.
A single precise "you appear 42% of the time"NoOne number hides a wide variance band; it will differ next run.
Exact app wording or citations, from API samplingNoAPI and consumer app share only about 24% of brand mentions.
"Your rank position in AI"NoSame-order odds are roughly 1 in 1,000; a rank is close to random.
Absolute traffic or "how many people saw it"NoNo tool sees how many humans actually received the answer.

Notice the pattern. Everything in the reliable column is relative and trended. Everything in the unreliable column is an absolute snapshot. That is the whole game. A tracker is a trend instrument, not a measuring tape.

How to read the numbers sensibly

Read a tracker the way you would read a stock chart: the slope matters, the single tick does not. Four habits keep you honest. Trust direction over absolutes, insist on multiple samples, look per engine rather than at a blended score, and judge changes over weeks instead of days.

Used this way, a tracker answers the questions that matter: are we trending up, are we gaining on the competitor we care about, and which engine should we work on next. Those are relative, multi-sample, per-engine questions, and they are exactly the ones the data can support. For the full workflow, see our guide on tracking brand mentions across ChatGPT, Claude and Perplexity.

Questions to ask any AI visibility vendor

Before you trust a dashboard, make the vendor answer four questions. Their willingness to answer plainly tells you as much as the answers themselves. A tool that treats its methodology as a secret is asking you to trust a number you cannot audit.

If you are still comparing options, our roundup of the best GEO tools for SaaS in 2026 covers how the major trackers answer these, and if the whole category is new to you, start with what generative engine optimization actually is.

Where GEOHATS hits exactly the same limits

It would be dishonest to critique every other tracker and pretend ours escapes physics. It does not. GEOHATS samples AI engines through their APIs, the same as most of the category, so we inherit the same API-vs-app gap and the same run-to-run variance the research describes. We built the product around that reality rather than hiding it.

That is why our dashboard leads with trends and share-of-voice movement, not a single hero percentage, and why the score carries a "provisional" label when it is based on only one sample per engine. It is also why we do not sell an "AI rank position," a metric the SparkToro data shows cannot be measured reliably. We would rather show you a defensible direction than a precise-looking number we cannot stand behind.

Where GEOHATS differs is what happens after measurement: instead of stopping at a score, the tool writes the page built to win the citation and helps you publish it, then watches whether your mention rate actually moves. The number is the starting gun, not the finish line. You can see how the pricing works on our pricing page.

Frequently asked questions

What determines ai visibility tracker accuracy?
Two things dominate ai visibility tracker accuracy: how many times the tool samples each prompt, and whether it queries the model API or the consumer app. More samples turn noisy single answers into a stable rate, the way more poll respondents shrink a margin of error. And because the API and the app share only about 24% of brand mentions (Surfer, via Gumshoe, 2026), a tool's data source decides whether its numbers reflect what your buyers actually see. Everything else, the pretty score included, sits on top of those two choices.
Do the numbers in Profound or Peec AI mean anything at all?
Yes, as long as you read them as trends rather than absolutes. The direction of your mention rate over weeks, your share of voice against a named competitor, and which engines cite you are all defensible signals. What is not defensible is treating a single "you appear 42% of the time" figure, or an "AI rank," as ground truth. Both Profound and Peec are measuring a probabilistic system, so their value is in relative movement over time, not in any one snapshot.
Why do two AI visibility tools give me different scores?
Usually because they sample differently. One may query the API while another scrapes the app, one may run each prompt 50 times while another runs it 5, and they may weight engines differently. Given that even one model's fast and reasoning modes overlap on only 25.6% of cited sources (Semrush, 2026), small methodology differences produce large score differences. Disagreement between tools is expected. Pick one method, stay consistent, and watch its trend rather than chasing a "true" absolute number that does not exist.
If AI answers are this random, is tracking even worth it?
It is, because randomness at the single-answer level averages into stable patterns at scale. SparkToro found that while exact lists almost never repeat, the top brands in a category still showed up in 55 to 77% of responses (SparkToro, 2026). That consistent presence is real and trackable. The mistake is not tracking; it is over-reading one run. Sample enough, trend over weeks, and the signal separates cleanly from the noise.
How many samples does a tracker need before I trust the number?
There is no magic threshold, but more is always better and one is close to useless. Thinking Machines Lab got 80 different outputs from 1,000 runs of a single prompt (Thinking Machines Lab, 2025), so a handful of samples still leaves a wide error band. Treat any rate built on one or two checks per engine as a rough presence indicator only. Rates built on dozens of runs per engine, tracked over weeks, are the ones worth putting in a report.
Can a tracker tell me my exact rank in ChatGPT?
No, and be wary of any tool that claims it can. SparkToro measured roughly 1 in 1,000 odds of two runs returning brands in the same order (SparkToro, 2026), which means an "AI rank position" is close to random from one run to the next. Presence and rough frequency are measurable. A precise ordinal rank is not. A vendor selling a rank number is selling precision the underlying system does not have.

Measure the trend, then fix it

GEOHATS tracks your mention rate across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews, then writes and publishes the page built to win the citation. Trends you can trust, plus the fix.

See how it works

Related reading

More answers