How to check whether AI is recommending your brand
Open ChatGPT, Perplexity, Gemini and Microsoft Copilot in fresh sessions, logged out if you can manage it. A logged-in account carries memory of past conversations that can quietly shape the answer: OpenAI's own Memory FAQ confirms that saved details from earlier chats are "always considered in future responses" unless you delete them. Ask the questions a real customer would type: "best [category] in [your city]", "who does [service] well", "recommend a [your industry] company". Note whether your brand appears, where it sits in the list, and whether the tool names a page it pulled the answer from.
Then do it again. Run the same five or six prompts tomorrow, and again next week, across a different device or account if you have one. Write the results down somewhere you'll actually look at again; a spreadsheet is fine. One afternoon of this costs nothing and beats guessing. It will not, on its own, give you a reliable score. The next section explains why.
Why your first check is almost meaningless
Rand Fishkin of SparkToro, working with Patrick O'Donnell of Gumshoe.ai, published the clearest answer to that question in a study of AI brand-recommendation consistency on 28 January 2026, updated 28 June 2026. 600 volunteers ran 12 prompts through three AI tools, ChatGPT, Claude and Google's AI Overview and AI Mode, using their own default settings rather than a controlled lab, to capture the variety a real user actually sees. That produced 2,961 total responses, covering mostly consumer product categories with a smaller number of service and local categories mixed in.
The finding: less than a 1 in 100 chance that ChatGPT or Google's AI, asked the same prompt 100 times, returns the same list of brands twice. For the exact same list in the exact same order, the odds fall to close to 1 in 1,000. Personalisation plays a part, per the memory mechanism above, and so does the tools' own run-to-run behaviour, which is precisely what the 2,961 runs were built to measure. A brand can appear in one run and be absent from the next for reasons that have nothing to do with how good the brand is.
What the SparkToro study found
What gets left out of most summaries of this study is what its author concluded from it. He changed his mind partway through the research. Rank position, read from a single run, is noise. But visibility measured across dozens to hundreds of prompts, not one, is a reasonable metric to track. The study is an argument against trusting a single prompt. It is not an argument against measuring the pattern, and treating it as one misreads what its own author concluded.
That distinction separates a useless number from a useful one. A single "am I mentioned" check tells you what happened in that one run and nothing more. The same check, repeated across a fixed set of prompts and tracked over weeks, starts to show something closer to a real signal: a share of runs where the brand appears, rising or falling over time, rather than a single yes or no.
What you can measure reliably, and what you cannot
Two further pieces of research point at the same underlying problem: being retrieved, or even cited, is not the same as being named.
Semrush, working with SEO analyst Kevin Indig, logged 3,981 domain appearances across 115 prompts in 14 countries, across four AI engines: ChatGPT, Google AI Overviews, Gemini and Google AI Mode (Perplexity and Copilot weren't covered). Its ghost citations study, published 11 September 2025, found 61.7% of citations were "ghost citations": the AI used the page as a source link but never named the brand in the answer text. Only 38.3% of appearances included the brand name at all.
Separately, AirOps tracked 548,534 pages ChatGPT retrieved while generating answers to 15,000 queries. Its report on retrieval and citation, dated 12 March 2026, found only 15% of retrieved pages made it into the final response; the rest were pulled up and discarded. That figure covers ChatGPT only, with no country or language scope stated.
Put together, this is what's reliably measurable: whether a brand or page is retrieved at all, whether it's cited as a source, and separately, whether it's actually named in the text a user reads. Those are three different outcomes. A tool that reports only one of them, and calls it "visibility", isn't telling the whole story. Whether an algorithm favours you over a named competitor on any given day is not something this research, or any tool built on it, claims to measure reliably.
The tools, and what each one really measures
A handful of paid products automate the check above, run at a scale no person could manage by hand.
Profound queries the consumer-facing apps directly, in its own words "the front-end experiences that normal users see", across nine named platforms: ChatGPT, Perplexity, Claude, Microsoft Copilot, Google AI Overviews, Google AI Mode, Gemini, Grok and DeepSeek. It reports visibility scores, share of voice, sentiment, citation sources and competitor rankings. Pricing isn't published on its site; it's sold direct to enterprise buyers.
Otterly.AI starts at $29 a month for its Lite tier: 15 tracked prompts across four engines (ChatGPT, Google AI Overviews, Perplexity, Microsoft Copilot), with Claude, Gemini and Google AI Mode as paid add-ons.
Ahrefs Brand Radar tracks mentions across six named AI platforms and separates the "mentioned" report from the "cited" report, the same distinction the Semrush research above found matters. It's now sold as a standalone tool as well as inside an Ahrefs subscription; check current pricing on Ahrefs' own site rather than taking a figure from any article, this one included.
None of these tools samples every possible query a customer might type. Each works from its own fixed set of prompts, run on a schedule: the manual method above, automated and scaled up.
What Google offers, and what it does not
Google is the only major AI or search platform to have shipped anything resembling a visibility report for its own AI answers, built into Search Console. It launched new Search Generative AI performance reports on 3 June 2026, and its own announcement is explicit that the rollout is limited: "we're rolling these reports out to a subset of websites, allowing us to thoroughly test them and receive feedback before making them widely available." Not every site has access yet.
Where it is available, it covers four things: impressions (how often pages appeared in AI features), which pages appeared, which countries, and which devices. That's the complete list Google names. There's no query-level breakdown showing which prompt triggered which appearance, and no click data specific to an AI answer.
Outside Google, nothing comparable exists as something a typical business can simply sign up for. Perplexity's only analytics offering, its Publishers Program, is invite-only, launched with a small named list of partners: TIME, Der Spiegel, Fortune, Entrepreneur, The Texas Tribune and WordPress.com. Anthropic publishes an Enterprise Analytics API, but it reports an organisation's own usage of Claude, cost and seat adoption, not how Claude describes a brand to other users. Neither substitutes for the check this article is about.
Turning one reading into a pattern
A single reading, whether it comes from an afternoon of manual prompts or a paid tool's dashboard, is not a result. It's one data point. The value starts once you have several, taken the same way, over time.
Track the same fixed set of prompts on a schedule, weekly or monthly, and watch the trend rather than any single figure: is the brand appearing more often across the same questions than it was last month, or less? Is the balance between "mentioned" and "cited" shifting, per the distinction covered above? That's the version of tracking the research actually supports: not a score checked once and filed away, but a repeated measurement that gets more meaningful the longer it runs.
If the trend is flat or falling, the fix usually sits upstream of any tracking tool, in whether an AI system can read and understand the site at all. We've covered how AI Overviews choose which brands to cite and what makes a website readable to AI separately, along with what an llms.txt file does and doesn't do. For the playbook on actively earning citations from ChatGPT specifically, rather than just measuring where a brand currently stands, that's covered on its own too, so it isn't repeated here.
A single reading, free or paid, is not a verdict on whether AI recommends a brand. Don't let one afternoon's check, or one paid dashboard's snapshot, decide anything on its own: run the same prompts again next week before treating this week's number as real.
↳ Frequently asked
01Is AI visibility tracking worth paying for?
Depends on scale. The free manual method costs nothing but time, and covers a handful of prompts checked repeatedly across a few tools. Paid tools automate that across more platforms and prompts than most people would run by hand, which matters more once tracking a whole market rather than a handful of terms.
02What is answer engine optimization (AEO)?
The practice of making sure AI tools can find, understand and cite a site's content when generating an answer, the AI-era counterpart to traditional SEO. Covered in full in what is answer engine optimization.
03Why did my brand appear in ChatGPT yesterday and not today?
Most likely the run-to-run variation covered above, not something that changed on the site overnight. The SparkToro study found under a 1 in 100 chance of an identical brand list twice for the same prompt, so one day's absence tells you very little by itself.
04Does Google's Search Console AI report show everything I need?
No. Where available, still a limited rollout as of its June 2026 launch, it shows impressions, pages, countries and devices, nothing else: no query-level detail, no click data specific to AI answers. It covers only Google's own AI surfaces, not ChatGPT, Perplexity, Claude or Gemini's chat product.
05How many prompts do I need before the number means something?
The research sets no fixed threshold, but its own framing is "dozens to hundreds" of prompts, not one. Paid tools' entry tiers give a sense of the working range: Otterly's cheapest plan tracks 15 prompts; Ahrefs Brand Radar covers six platforms per brand. More prompts, tracked consistently over time, produce a steadier picture than fewer.