GEOHATS Blog

GEO · How-to · 2026

How to Track Brand Mentions in AI Answers (ChatGPT, Claude, Perplexity) in 2026

The manual method first: a prompt-set spreadsheet you can start today. Then the tools, with verified pricing, for the day the spreadsheet stops scaling.

By the GEOHATS team · Last updated: August 4, 2026

TL;DR · Quick answer

Pick 20 to 30 questions your buyers actually ask, run each one in ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews on the same day every week, and log four fields per answer: mentioned, cited, position, competitors named. A spreadsheet handles one brand at small scale; dedicated tools automate it from $29 per month.

On this page

  1. Why a single check misleads you
  2. The DIY method, step by step
  3. Engine differences that change what you measure
  4. The honest measurement caveats
  5. Where DIY breaks down
  6. Tools that automate it (verified pricing)
  7. The four metrics that matter
  8. FAQ

Buyers now ask ChatGPT, Claude and Perplexity the questions they used to type into Google. Gartner predicted in February 2024 that traditional search engine volume would drop 25% by 2026 as AI assistants absorb those queries (Gartner, Press Release 2024). The only way to know whether those assistants name you is to measure it, repeatedly.

87%
of SearchGPT citations matched Bing's top organic results in a 500+ citation study (Seer Interactive, 2025)
4%
source overlap between ChatGPT's web app and the OpenAI API across 1,000 prompt runs (Surfer, 2026)
6.49% → 25%
swing in how often Google showed AI Overviews during 2025, across 10M tracked keywords (Semrush, 2025)

Why a single check misleads you

Ask an AI engine the same question twice and you can get two different answers, with different brands and different sources. That is not a bug in your test; it is how these systems work. One check tells you almost nothing. A trend across repeated, scheduled checks tells you everything.

The cause is documented: LLM inference is nondeterministic in production. Research from Thinking Machines Lab shows that even at temperature zero, server load and batch-size variation alone produce different outputs for identical prompts (Thinking Machines Lab, Defeating Nondeterminism in LLM Inference 2025).

The engines themselves also move constantly. Semrush tracked over 10 million keywords through 2025 and watched AI Overviews' presence climb from 6.49% of tracked queries in January to roughly 25% in July, then fall back to 15.69% by November (Semrush, AI Overviews Study 2025). A brand that "lost" its AI visibility in August may simply be caught in one of these swings.

The core principleThe unit of measurement is not "does ChatGPT mention us?" It is "in how many of our sampled answers this month did ChatGPT mention us, and is that share rising or falling?" Everything below follows from that.

How to track brand mentions in AI answers by hand (the DIY method)

The manual method needs a spreadsheet, one hour per week, and discipline. Run it for four weeks and you will know your mention rate, your citation rate, and which competitors AI engines prefer, before spending anything on tooling.

Step 1: Build a fixed prompt set

Write 20 to 30 questions your ideal buyer would ask an AI assistant. Not keywords, questions: "best invoicing software for freelancers in Germany", "[competitor] alternatives". Freeze the list. If the questions change every week, the trend line means nothing.

Step 2: Pick your engines

Cover the assistants your buyers use: ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews (search your question on Google and record what the AI box says). Use fresh chats and a clean browser profile to reduce personalization.

Step 3: Run on a schedule, and sample

Same day, every week. Because answers vary run to run, ask each question 2 to 3 times per engine and record each run as its own row. Sampling is what turns noise into a measurement.

Step 4: Record the same fields every time

This is the sheet. Nine columns, one row per answer:

The DIY tracking sheet: what to record for every AI answer
ColumnWhat to recordWhy it matters
DateRun dateEnables the trend line
EngineChatGPT / Claude / Perplexity / Gemini / AI OverviewsVisibility differs sharply per engine
QuestionThe exact prompt, verbatimKeeps runs comparable
Mentioned?Y/N: your brand named in the answer textYour headline metric: mention rate
Cited?Y/N: your domain linked as a sourceCitations drive clicks and trust
Position1st, 2nd, 3rd... brand named, or absentFirst-named brands win the recommendation
Competitors namedEvery rival brand in the answerFeeds share of voice
Sources citedDomains the engine linkedShows which pages to beat or get on
NotesSentiment, wrong facts, quoted phrasesCatches misinformation early

Step 5: Score it monthly

At month end, compute four numbers per engine: mention rate (answers naming you ÷ total answers), citation rate (answers linking you ÷ total), share of voice (your mentions ÷ all brand mentions), and the delta versus last month. Those four numbers are the report. If you then want the pages that close the gaps, our playbook on how to get cited by ChatGPT covers the content side.

Engine differences that change what you measure

The five engines are not five copies of the same thing. Each retrieves and cites differently, which changes both how you check and how you improve. Treat each engine as its own channel with its own scoreboard.

How the major engines source and cite answers (2026)
EngineRetrieval behaviorWhat to watch when tracking
ChatGPTWeb answers lean heavily on Bing's indexYour Bing rankings; citations often mirror Bing top 10
PerplexitySearch-first, numbered citations on nearly every answerCitation rate matters as much as mention rate; freshness weighs heavily
Google AI OverviewsAI summary above classic results, presence fluctuates by query typeWhether the box appears at all for your questions, then who it links
ClaudeOptional web search; often answers from model knowledgeMentions without citations are common; record both separately
GeminiGoogle-grounded answersOverlap with your Google rankings, which does not guarantee an AI mention

The ChatGPT row is the best-studied: Seer Interactive matched 500+ SearchGPT citations against search rankings and found 87% matched Bing's top organic results, against 56% for Google (Seer Interactive, SearchGPT Citations Study 2025). If ChatGPT ignores you, check Bing first; we broke down that failure mode in You rank #1 on Google but ChatGPT ignores you.

Perplexity attaches numbered, clickable citations to answers by design (Perplexity, Help Center 2026), so "cited" is unambiguous there in a way it is not in a casual ChatGPT reply.

The honest measurement caveats

Every method on this page, manual or tooled, has three limits. Vendors rarely say this part out loud, so we will.

1. API answers are not app answers. Most tracking tools query provider APIs, but the consumer apps wrap the same models in system prompts, memory and interface logic. Surfer ran 1,000 prompt executions comparing ChatGPT's web app against the OpenAI API and found only 24% brand overlap and 4% source overlap; the same test on Perplexity found 8% source overlap (Surfer, Scraped vs API Study 2026). API tracking is still a consistent, comparable signal, but it is a proxy for the consumer experience, not a replica. This applies to GEOHATS too: our scans query provider APIs.

2. One sample is an anecdote. As covered above, identical prompts produce different answers. Any tracker (or spreadsheet) that reports a score from a single run per question is reporting noise. Look for sampling, or do it yourself.

3. Personalization skews spot checks. Logged-in accounts carry history and location context, and your own ChatGPT telling you about your own brand proves nothing. Use clean sessions, or accept a tool's neutral API vantage point as the more comparable baseline.

Where DIY breaks down

The spreadsheet method fails on arithmetic, not on principle. Thirty questions across five engines is 150 answers per week. Add the 2 to 3 samples per question that nondeterminism demands and you are reading, classifying and logging 300 to 450 AI answers every week, roughly 1,500 or more rows a month, by hand.

At one to two minutes per answer that is a 5 to 10 hour weekly job, forever. Most teams quietly stop by week six. That is what the paid tools actually sell: the discipline, automated.

Tools that track brand mentions in AI, with verified pricing

Five options worth shortlisting in 2026, from cheapest entry to enterprise. Prices checked August 4, 2026, against vendor pages or dated reviews as noted; confirm before buying, this market reprices often. A longer teardown lives in our best GEO tools for SaaS guide.

AI brand-mention trackers compared (entry tier, checked August 2026)
ToolEntry priceEngines at entryPrompts at entryPricing model
Otterly.AI$29/mo (Lite)ChatGPT, AI Overviews, Perplexity, Copilot15Subscription
GEOHATS$49 one-time (lifetime)ChatGPT, Claude, Perplexity, Gemini, Grok (+AIO add-on)Per-site question setsLifetime license + your own AI keys
Peec AI€89/mo (Starter)ChatGPT, Perplexity, AI Overviews25Subscription
Profound$99/mo (Starter)ChatGPT only50Subscription
Trakkr$100/mo (Growth, 14-day trial)8 engines incl. Claude, Grok, DeepSeek50Subscription

Otterly.AI $29/mo · verified Aug 2026

The cheapest real entry point. The Lite plan tracks 15 prompts daily across ChatGPT, Google AI Overviews, Perplexity and Microsoft Copilot (Otterly.AI, Pricing 2026). Claude and Gemini cost extra as add-ons, and 15 prompts is tight. Best first paid step up from the spreadsheet.

Peec AI from €89/mo · review-verified 2026

A polished mid-market tracker with daily runs, position and sentiment metrics, and unlimited seats. The vendor publishes plans at peec.ai/pricing; 2026 reviews list Starter at €89/month for 25 prompts (ColdIQ, Peec AI Review 2026). Strong dashboards, but extra engines cost more, and it only measures; it does not produce the content that fixes the gaps it finds.

Profound $99/mo · verified Aug 2026

The enterprise reference brand in this category, with deep answer-volume analytics. Note the entry math: the $99 Starter tracks ChatGPT only, with 50 prompts; Perplexity and AI Overviews arrive at the $399 Growth tier (Profound, Pricing 2026). Excellent if you have the budget and a team to act on the reports; oversized for a single founder.

Trakkr $100/mo · review-verified 2026

The broadest engine coverage of the group: eight models per plan, including Claude, Grok, DeepSeek and Meta AI. There is no permanent free tier; you get a 14-day full-access trial, then Growth at $100/month for one brand and 50 prompts (Am I Cited, Trakkr Review 2026). Its free side tools (an AI site grader, an llms.txt generator) cost nothing to try.

GEOHATS $49 lifetime · our product

GEOHATS tracks your buyer questions across ChatGPT, Claude, Perplexity, Gemini and Grok on a schedule, records mentions and citations per engine, computes share of voice against competitors, and then does the part trackers skip: it writes the pages targeting the questions you are losing and publishes them to WordPress. Pricing is a one-time $49 lifetime license for 3 sites (bring your own AI API keys; you pay providers at cost, typically cents per scan).

Where it fits
One-time price instead of a subscription; tracking and content generation in one loop; MCP server so Cursor or Claude Code can drive it.
Honest cons
Newer product with a shorter track record than Profound or Peec; BYOK setup takes a few more minutes than a pure SaaS; auto-publishing is WordPress-only today; API-based scans carry the same app-vs-API caveat as every tool above.

Disclosure: GEOHATS is our product. We have kept the comparison factual and linked every competitor's pricing; verify current features on each vendor's site.

What a good tracking setup measures

Whether you stay manual or buy a tool, judge the setup by whether it produces these four numbers per engine, per month. If it cannot, it is a demo, not a measurement system.

Set expectations on timing, too: content changes typically take 2 to 8 weeks to show up in AI answers, because engines re-crawl and re-ground on their own schedules. Track weekly, judge monthly.

FAQ

What is the best way to track brand mentions in AI for free?

Use the spreadsheet method above: freeze 20 to 30 buyer questions, run them weekly in clean sessions across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews, and log mentioned, cited, position and competitors named for every answer. It costs an hour or two a week and produces a real mention-rate and share-of-voice trend within a month.

How often should I check whether ChatGPT or Perplexity mentions my brand?

Weekly, on the same day, with the same frozen question set. Daily checking adds noise faster than insight; monthly checking cannot separate a real trend from the platform swings Semrush documented across 2025. Weekly sampling with monthly scoring is the practical middle.

Why does ChatGPT give a different answer every time I ask the same question?

LLM inference is nondeterministic: sampling temperature plus server-side batch variation means identical prompts legitimately produce different outputs, as Thinking Machines Lab's research shows even at temperature zero. For tracking, that means one answer is an anecdote. Ask each question multiple times, record every run, and report the rate at which you appear, not whether you appeared once.

Do API-based tracking tools show what real users actually see?

Not exactly, and any honest vendor should say so. Surfer's 1,000-run comparison found only 24% brand overlap and 4% source overlap between ChatGPT's web app and the OpenAI API. API tracking is still valuable because it is consistent over time, which is what a trend needs. Treat absolute numbers as a proxy and the trend direction as the finding. GEOHATS scans work this way too.

My brand ranks #1 on Google. Doesn't that mean AI engines mention me?

No. ChatGPT's web citations correlate with Bing, not Google: Seer Interactive found 87% of SearchGPT citations matched Bing's top organic results, versus 56% for Google. Claude often answers from model knowledge without searching at all. Track each engine separately and fix each one on its own terms.

Which metrics should a brand-mention tracker report?

Four per engine: mention rate, citation rate, share of voice against named competitors, and the week-over-week trend on a frozen prompt set. Position (are you the first brand named?) is a strong fifth. If a tool reports a single "AI visibility score" without exposing these components and its sampling, you cannot audit what moved or why.

Stop checking by hand. Start closing the gaps.

GEOHATS tracks your buyer questions across five AI engines on your own API keys, then writes the pages that win the ones you are losing. One-time $49 lifetime license, no subscription.

See lifetime pricing

Related reading

More answers