First, check whether it's actually happening
Before you touch your website, your content, or your budget, run the test the study itself points to. Open a fresh chat (or an incognito window, so saved memory and past conversations don't shape the answer) and ask the exact question your prospect would ask, the one where your competitor showed up and you didn't. Ask it again the next day. Ask it a third time a few days after that, ideally from a different device or account. Write down which brands appear each time, not just which one appears first.
You're checking two separate things. First, does your competitor show up consistently, in most of your attempts, or was it one appearance out of one try. Second, do you show up at all, anywhere in the list, even if not first. A brand that appears in most runs but rarely tops the list is a genuinely different situation from a brand that never appears. The SparkToro researchers found exactly this pattern when they tracked one hospital brand through ChatGPT: it showed up in 69 of 71 responses to a cancer-care prompt, a 97% visibility rate, but was only the top-named answer in 25 of those 71. High presence and low rank are not the same problem, and they don't call for the same response.
What the research found about consistency
The instability isn't a fluke of one prompt. Across the full SparkToro study, ordering was even less stable than list membership: roughly 1 in 1,000 runs produced the same brands in the same order twice. Run the same question through ChatGPT, Claude, or Google's AI Overview and AI Mode a hundred times, and you would see a hundred slightly different answers far more often than not.
It matters what kind of category you're in. The same study found that categories with fewer genuine alternatives produced far more consistent answers than categories with many. Cloud computing providers for SaaS startups, a market with a handful of credible options, showed high agreement across runs. Recently published science fiction novels, a market with hundreds of plausible answers, showed low agreement and wide swings in ranking. A small, specialist market behaves more like the first case than the second, which is worth knowing before you assume your result is typical of the study's headline number.
A separate signal worth naming directly: a Washington State University study reported in March 2026 found ChatGPT gave the same verdict only about 73% of the time when the identical prompt was repeated ten times. That study tested true or false judgments on business-research hypotheses, not brand recommendations, so it isn't a number to quote about your market. It is evidence that the underlying behaviour, the same question producing different answers, isn't unique to one type of prompt.
Not everyone agrees on what this means for tracking. At least one AI-visibility monitoring vendor has argued that brand mentions hold steadier than the wording around them, even when full answers vary. That claim has no published sample size or method behind it, and it sits against the more rigorous, more recent SparkToro data. We're not going to pretend that disagreement doesn't exist. What we will say plainly, because the same SparkToro research makes the point directly: rank position on any single run is close to noise, but visibility measured across dozens or hundreds of runs is a reasonable thing to track. Those are two different claims. Only the first one is unreliable.
Telling a real gap from noise
Sometimes the repeat test comes back the other way. Your competitor appears in most runs, you appear in few or none, and that pattern holds over days and different phrasings. That is a real signal, not sampling noise, and it deserves a real response.
This is exactly the distinction the SparkToro data draws out. A brand appearing in most runs, even without ever topping the list, behaves nothing like a brand that never appears. If you've run the repeat test three or four times over a week and your name simply doesn't come up while a named rival's does, in a category where there aren't dozens of credible alternatives to rotate through, that consistency is the finding. It stops being a question of luck and starts being a question of whether the engine can find and trust your business at all.
The same study offers one useful diagnostic here. When Google's AI was asked to recommend digital marketing consultants with e-commerce expertise, one competitor appeared in 85 of 95 responses, a single test, one prompt, one engine, but directly analogous to the kind of question a prospect might ask about your category. A result that lopsided in a market with a limited set of credible answers is a genuine gap in how consistently and how clearly an AI system can find and trust a brand, not an artefact of asking once.
Neither OpenAI nor Google publishes the full mechanics of how an answer gets assembled. What they've confirmed is narrower: OpenAI's own documentation on ChatGPT search says it may use IP address to estimate location for results like restaurants and local news, and its memory documentation says memory personalises "your experience" generally. Neither statement confirms that location or memory changes which non-local business gets named in a recommendation-style answer. That link is a reasonable guess, not a documented fact, and we'd rather say that than dress a guess up as certainty.
Fixing a real gap
If the repeat test shows a real, consistent gap, the fix isn't a mystery and we've already written it up in full. Two things drive whether an AI engine can find, trust, and name your brand: whether your site is structured so the engine can actually read and cite it, and whether your business is corroborated by other sources the engine already trusts, not just your own pages. We cover the technical side, structure, page speed, the specific fixes to run first, in Make Your Website Readable to AI. We cover the trust and citation side, what actually gets a brand named, in How to Get Cited by ChatGPT. Both are the whole method, not a teaser.
If the repeat test instead shows the sighting was one draw from a rotating list, the honest answer is that there's nothing to fix yet. Keep running the test on a schedule, watch whether the pattern holds, and don't spend money reacting to a single screenshot.
Guarantees no one can back
No engine publishes its selection algorithm, and nobody outside the AI labs can hand you a formula that works across categories, engines, and time. Be wary of anyone who claims otherwise. A guaranteed ChatGPT citation, a promised ranking position, an "AI SEO score" pitched as a settled science: none of that is backed by anything either company has confirmed, and the honest limits of what's known are narrower than most sales pages admit.
Chasing rank position on a single run is the clearest way to waste a budget, since the research above found that ordering is close to random on any one attempt. And treating your own IP-based location or a competitor's saved memory as an explanation for a single result is guessing at mechanics neither company has confirmed. If you want the fuller picture of where SEO and AEO budget actually pays off, SEO vs AEO: Where to Invest First and How AI Overviews Choose Which Brands to Cite cover that ground properly.
Run the test three times before deciding anything; a pattern that holds across a week is a different story to one screenshot. If it holds, tell us about your business, and we'll help you work out what's actually happening in your category before you spend anything fixing it.
↳ Frequently asked
01Does asking ChatGPT the same question twice count as a real test?
It's a start, but the SparkToro data suggests three to five repeats, on different days, gets you closer to a reliable read than two. A single repeat rules out the most obvious false alarm; it doesn't confirm a stable pattern either way.
02Why would ChatGPT mention a competitor by name and never mention us at all, even once?
That could be a genuine gap in corroboration or site structure, or it could be a category with many credible answers and a naturally low chance of naming any one brand on a given run. The repeat test over several days is what tells those two apart.
03If we never show up, does that mean ChatGPT doesn't know we exist?
Not necessarily. Absence in a handful of manual runs isn't proof of absence overall, especially in a category with many possible answers. It's a reason to test properly before concluding anything.
04Should we pay for a service that promises we'll be recommended by ChatGPT?
Be sceptical of any fixed guarantee, since neither OpenAI nor Google has published the mechanics a vendor would need to guarantee that outcome. A service that improves how readable and corroborated your site is can genuinely help; a service that promises a named outcome is promising something nobody can currently deliver on demand.
05Does this apply the same way to Google's AI Overviews as it does to ChatGPT?
The instability finding covers both; the SparkToro study ran prompts across ChatGPT, Claude, and Google's AI Overview and AI Mode. The mechanics behind each differ, which is why we've written about Google's citation behaviour separately.