GEO measurement
How Accurate Are AI Visibility Trackers? What the Numbers Really Mean
Tools like Profound and Peec AI put a confident number on your screen. Here is how much of that number you can actually trust, and how to read it without fooling yourself.
Quick answer
AI visibility trackers are directionally useful, not precise gospel. Because language models return different answers to the same question run to run, any single "you appear 42% of the time" figure is noisy. What you can trust is the direction of your mention rate over many samples, per engine, over weeks. Treat absolute numbers with skepticism.
Every AI visibility tool sells the same promise: a dashboard that tells you how often ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews mention your brand. The dashboards look precise. A score, a percentage, a rank. The uncomfortable truth is that the underlying system is probabilistic, and precision on top of noise is still noise.
This is not a reason to throw the tools out. It is a reason to read them correctly. Below is what the research actually shows about how repeatable AI answers are, why two credible tools can disagree, and which parts of a tracker report you can take to a client meeting.
<1 in 100
chance two identical runs return the same brand list, across 2,961 prompts on ChatGPT, Claude and Google AI (SparkToro, 2026)
80
distinct outputs from 1,000 runs of one prompt at temperature 0 (Thinking Machines Lab, 2025)
24%
brand overlap between ChatGPT's API and its consumer app over 1,000 prompts (Surfer, via Gumshoe, 2026)
Why the numbers move even when nothing changed
AI answers are nondeterministic by design and, it turns out, by accident too. Run the same prompt twice and you can get a different set of brands, a different order, and different sources. This is the root cause of every accuracy question about trackers.
The scale of the variance is larger than most marketers assume. SparkToro's Rand Fishkin and Gumshoe's Patrick O'Donnell ran 2,961 prompts across ChatGPT, Claude and Google's AI Overviews, repeating each prompt 60 to 100 times. They found less than a 1 in 100 chance that any two runs return the same brand recommendations, and roughly 1 in 1,000 odds of the same list in the same order (SparkToro, January 2026). Fishkin's conclusion was blunt: any tool that reports a "ranking position in AI" is selling a metric the data cannot support (Search Engine Journal, 2026).
You might expect this to disappear at temperature 0, the setting meant to make a model deterministic. It does not. Thinking Machines Lab sampled 1,000 completions of a single prompt at temperature 0 and got 80 unique outputs, with the very first divergence appearing at token 103 (Thinking Machines Lab, September 2025). The cause is not the prompt. It is batch-size dependence in the inference kernels: the same request gets grouped differently depending on server load, and the math shifts slightly each time. For a tracker sampling a live production endpoint, this variance is baked in.
The takeaway
A single sample is a coin flip dressed up as a measurement. This is why serious trackers sample the same prompt many times and report a rate, not a one-shot result.
The API-vs-app problem, the caveat nobody puts on the label
The biggest accuracy gap is not run-to-run noise. It is that most trackers query models through their API, while your buyers use the consumer app, and the two return meaningfully different answers. Any tool measuring the API is a proxy for what customers actually see, not a mirror of it.
Surfer compared 1,000 ChatGPT prompts sent through the API against the same prompts scraped from the ChatGPT web app. The brand overlap between the two was 24%, and the API returned citations only about 75% of the time while the app always did (Surfer, reported by Gumshoe, 2026). The consumer app wraps the model in system instructions, retrieval, and formatting rules that the raw API never sees.
Vendors handle this differently, and it is worth knowing which method your tool uses before you trust its numbers. Broadly, some platforms sample the API for depth and volume, while others scrape or simulate the app interface for fidelity to what users see; neither approach is "correct," they answer different questions (Brainlabs, 2026). Check Profound and Peec AI directly for their current methodology rather than trusting a comparison blog, including ours. If a tracker will not tell you whether its data comes from an API or the app, that silence is the answer.
The engines themselves also disagree with each other, which is why a per-engine breakdown matters more than a blended score. An analysis of 680 million AI citations found only about 11% domain overlap between the sources ChatGPT and Perplexity cite (Averi, 2026). Even within one model, Semrush found only 25.6% of cited domains overlapped between ChatGPT's fast and reasoning modes (Semrush, June 2026). Growth advisor Kevin Indig summed it up in that study: "The brand that wins under minimal reasoning is not the brand that wins under high reasoning. These are two different systems."
AI visibility tracker accuracy: what the numbers measure
AI visibility tracker accuracy comes down to one distinction: a mention rate over many samples is a real, defensible statistic, while a single "did it mention me" check is noise. Trackers earn their value by sampling enough to turn the coin flips into a stable percentage you can trend, not by pretending any single answer is the truth.
Think of it like polling. No pollster claims one phone call predicts an election. They call thousands of people and report a range with a margin of error. A good tracker does the same thing with prompts: it runs each buyer question many times per engine and reports the share of runs that mentioned you. The more samples, the tighter the estimate. This is why sample count is the single most important number a vendor can disclose, and the one they most often hide.
What this means in practice: a mention rate that moves from 20% to 35% over a month of consistent sampling is a real signal worth acting on. A jump from 40% to 44% inside a single week, or a metric built on one check per engine, is likely inside the noise band. The number is only as accurate as the sampling behind it.
What trackers can and cannot tell you reliably
The honest way to use a tracker is to trust it for relative movement and treat its absolute figures as estimates. The table below splits the typical dashboard into signals you can defend and signals you cannot, based on the variance research above.
| Signal on the dashboard | Reliable? | Why |
|---|---|---|
| Your mention-rate trend over weeks | Yes | Many samples average out the run-to-run noise; direction is stable. |
| Share-of-voice movement vs named competitors | Yes | A relative comparison inside one sampling method is apples to apples. |
| Which engines mention you at all | Yes | Presence or absence is far more stable across samples than exact rate. |
| A single precise "you appear 42% of the time" | No | One number hides a wide variance band; it will differ next run. |
| Exact app wording or citations, from API sampling | No | API and consumer app share only about 24% of brand mentions. |
| "Your rank position in AI" | No | Same-order odds are roughly 1 in 1,000; a rank is close to random. |
| Absolute traffic or "how many people saw it" | No | No tool sees how many humans actually received the answer. |
Notice the pattern. Everything in the reliable column is relative and trended. Everything in the unreliable column is an absolute snapshot. That is the whole game. A tracker is a trend instrument, not a measuring tape.
How to read the numbers sensibly
Read a tracker the way you would read a stock chart: the slope matters, the single tick does not. Four habits keep you honest. Trust direction over absolutes, insist on multiple samples, look per engine rather than at a blended score, and judge changes over weeks instead of days.
- Trend, do not screenshot. One number this Tuesday means little. The line over eight weeks means everything. If a tool only shows today's figure, it is showing you a coin flip.
- Demand multiple samples per prompt. A rate built on 20 runs per engine is defensible; a rate built on one is not. The variance research makes single-sample checks close to worthless for anything but a rough presence check.
- Read per engine. With roughly 11% source overlap between ChatGPT and Perplexity (Averi, 2026), a blended cross-engine score can hide that you are strong in one place and invisible in another.
- Give it weeks. Citations move slowly. A change worth acting on shows up as a sustained shift, not a one-week wobble. Our own guidance to customers is that new pages take two to eight weeks to start getting cited.
Used this way, a tracker answers the questions that matter: are we trending up, are we gaining on the competitor we care about, and which engine should we work on next. Those are relative, multi-sample, per-engine questions, and they are exactly the ones the data can support. For the full workflow, see our guide on tracking brand mentions across ChatGPT, Claude and Perplexity.
Questions to ask any AI visibility vendor
Before you trust a dashboard, make the vendor answer four questions. Their willingness to answer plainly tells you as much as the answers themselves. A tool that treats its methodology as a secret is asking you to trust a number you cannot audit.
- API or app? Do you query the model API or the consumer interface? If API, how do you account for the roughly 24% overlap with what users actually see?
- How many samples? How many times do you run each prompt, per engine, per cycle? A rate without a sample count is decoration.
- Which engines, and reported separately? Which engines do you cover, and can I see each one on its own rather than a single blended figure?
- Is this a trend or an absolute? Does the headline number claim a real-world volume, or is it a directional estimate? The honest answer is almost always the second one.
If you are still comparing options, our roundup of the best GEO tools for SaaS in 2026 covers how the major trackers answer these, and if the whole category is new to you, start with what generative engine optimization actually is.
Where GEOHATS hits exactly the same limits
It would be dishonest to critique every other tracker and pretend ours escapes physics. It does not. GEOHATS samples AI engines through their APIs, the same as most of the category, so we inherit the same API-vs-app gap and the same run-to-run variance the research describes. We built the product around that reality rather than hiding it.
That is why our dashboard leads with trends and share-of-voice movement, not a single hero percentage, and why the score carries a "provisional" label when it is based on only one sample per engine. It is also why we do not sell an "AI rank position," a metric the SparkToro data shows cannot be measured reliably. We would rather show you a defensible direction than a precise-looking number we cannot stand behind.
Where GEOHATS differs is what happens after measurement: instead of stopping at a score, the tool writes the page built to win the citation and helps you publish it, then watches whether your mention rate actually moves. The number is the starting gun, not the finish line. You can see how the pricing works on our pricing page.