Guide

How to Track Whether ChatGPT Recommends You or a Competitor

We ran 60 buyer questions against four AI engines for a month. ChatGPT's cited sources changed 99.7% of the time between runs a day apart. Here is what to track.

Tarang AgarwalAugust 10, 20266 min read min read
Title card for an article on tracking whether ChatGPT recommends your brand or a competitor

You track it by asking the same buyer questions on a fixed schedule and recording two things each time: which brands get named, and which sources the answer cites. One check is not enough to know where you stand. Across 4,220 consecutive runs a median of 24.0 hours apart, ChatGPT came back citing a different set of sources 99.7% of the time.

That number needs care, though, because it is easy to misread in a way that sells monitoring harder than the evidence supports. Sources churning is not the same as your brand appearing and disappearing. Those two things move at very different speeds, and only one of them is a reason to look every day.

What follows is measured from our own tracking between 10 July and 10 August 2026: 130 buyer-intent prompts across four engines, 4,880 runs on 20 distinct days. That gives 520 prompt-and-engine series to compare. Sixty of those prompts are still live and 70 were archived partway through, for reasons that turn out to matter later in this piece. GetIntel publishes this and sells an AI visibility tracker, so treat the recommendation accordingly. The full dataset is here and every figure below recomputes from it.

What changes between one run and the next

Almost everything, on the citation side. We compared each pair of consecutive runs for the same prompt on the same engine and asked whether the set of cited domains was identical. It usually was not.

enginerun pairspairs where sources changedrate
ChatGPT1,0501,04799.7%
Gemini1,0581,04799.0%
Google AI Overviews1,03998795.0%
Perplexity1,07382076.4%
all four4,2203,90192.4%

The average overlap between two consecutive citation lists was 0.236 by Jaccard similarity. Two answers to the same question, a day apart, share under a quarter of their sources.

That 0.236 is pooled across all four engines and it is worth splitting, because the engines are nothing like each other. Perplexity shares 51.4% of its sources with itself a day later, Google AI Overviews 16.9%, ChatGPT 16.7% and Gemini 10.1%. Quoting the pooled figure as if it described ChatGPT would overstate ChatGPT's stability by more than a third.

One caveat on "a day apart". The median gap between consecutive runs was 24.0 hours and 73.4% fell in the 20 to 30 hour band, but the mean was 37.2 hours because a minority of gaps are much longer. We measured on 20 distinct days inside a 31-day window rather than all 31, so the figures describe the typical daily case and not every pair.

Perplexity is the outlier worth noticing. It repeated its exact source set 253 times out of 1,073, where ChatGPT managed it 3 times out of 1,050. Put the other way round, Perplexity is about 82 times more likely to hand you the same citation list twice running, which fits its design as a search-first product with a visible citation list rather than a generative answer that reaches for sources as it goes.

Does your brand's presence move as much as your sources?

No, and this is where the case for daily checking gets more interesting than the headline number suggests.

We counted how often our own brand flipped between cited and not-cited across all 4,360 consecutive pairs. It flipped 169 times. That is 3.9%. (The presence figures use 4,360 and the churn figures above use 4,220, because comparing two source lists needs at least one of the two runs to have returned sources.) So while the sources underneath an answer are almost completely reshuffled from one day to the next, the question of whether a given brand gets named is comparatively stable.

Both facts are true at once, and they are not in tension once you separate them. The engine is re-retrieving its evidence constantly. It is drawing broadly similar conclusions from it.

Anyone selling you daily monitoring on the strength of the 99.7% figure alone is quoting the volatile number and letting you assume it describes the stable one.

Bar chart comparing two rates per engine. Cited sources changed on 99.7% of consecutive ChatGPT runs, 99.0% Gemini, 95.0% Google AI Overviews and 76.4% Perplexity, while our brand's presence flipped on only 7.7%, 0.2%, 3.4% and 4.2% respectively.
Bar chart comparing two rates per engine. Cited sources changed on 99.7% of consecutive ChatGPT runs, 99.0% Gemini, 95.0% Google AI Overviews and 76.4% Perplexity, while our brand's presence flipped on only 7.7%, 0.2%, 3.4% and 4.2% respectively.

Presence moves at very different speeds depending on which engine you ask, which the aggregate hides:

enginerun pairstimes our presence flippedflip ratewe were cited in
ChatGPT1,090847.7%7.2% of runs
Perplexity1,090464.2%8.5% of runs
Google AI Overviews1,090373.4%2.9% of runs
Gemini1,09020.2%0.1% of runs

These use all 4,360 consecutive pairs. The churn figures above use 4,220, because comparing two source lists requires at least one of the runs to have returned sources at all.

Perplexity deserves its own note, because checking it once is the single most common way people conclude they are invisible. It cited us in 8.5% of runs, the highest of the four, while flipping presence on only 4.2% of consecutive pairs and on 17 of its 130 series. So a one-off Perplexity check that comes back without your name is usually telling the truth about that day and still tells you almost nothing about the trend. In our data 11 of those 17 unstable series were absent more often than present, which is exactly the situation where one check reads as a verdict and is really a sample of one.

Gemini sits at the opposite extreme. Two flips in 1,090 pairs is not stability worth celebrating when the underlying rate is 0.1%: it has essentially never named us at all, so there is nothing to flip.

That also gives you a working definition of gaining or losing ground. A trend line is moving when the share of runs naming you changes across a full week of the same prompts on the same engine. A single day crossing from named to unnamed is not a trend, it is one of the 169 flips above.

The case for daily tracking is narrower than it sounds

Here is the honest version. Over the month we measured 520 distinct prompt-and-engine pairs. Of those, 65 changed state at least once: cited on some days, absent on others. That is 12.5%, or about one in eight.

For the other seven in eight, a single spot-check would have told you the truth, and checking daily would have told you the same thing thirty times over.

The value sits entirely in that one-in-eight. On those 65 pairs, a single check landed on the minority state 23.2% of the time. Roughly one check in four on a genuinely unstable prompt gives you a reading that does not represent the month. You cannot tell in advance which of your prompts are the unstable ones, which is the actual argument for running them on a schedule rather than by hand when you remember.

So the reason to track daily is not that everything moves constantly. It is that a minority of your prompts move, you do not know which, and those are precisely the ones where a one-off check misleads you.

What to record on each run

Four fields cover it, and the second is the one people skip.

  • Whether you were named. The binary. This is the number that belongs on a trend line, because it is stable enough that a change in it means something happened.
  • Which competitors were named alongside you. Presence on its own tells you very little. Being named third of three is a different situation from being named alone, and only the competitor list distinguishes them. Our guide to finding which competitors ChatGPT recommends covers how to pull that list.
  • Which sources the answer cited. Expect this to be noisy, given the 92.4% churn. Do not put it on a daily trend line. Aggregate it over weeks instead and look at which domains recur, which is a far more stable signal than any single day's list. We wrote up what that aggregate looks like across engines.
  • The full answer text. Cheap to store and the only way to reconstruct why a number moved once it has moved.

Setting up a daily check

The mechanics matter less than the consistency. Run the same prompts, at roughly the same time, through an interface a buyer would actually use. That last part is not a detail: the API and the rendered interface share only 13.6% of their sources, less than the interface shares with itself a day apart.

Fix the prompt wording and leave it alone. Any edit to the question resets your history, because you are no longer measuring the same thing. We learned that the hard way. An unrelated profile edit re-minted our own prompt set on 3 August, archiving 70 prompts mid-measurement, and we wrote up what that cost us rather than quietly restating the numbers.

Run every engine you care about, not just ChatGPT. The 76.4% to 99.7% spread above means engine choice changes what you observe more than most other variables. Rolling four engines into a single visibility score is convenient for reporting and hides exactly this.

Set the alert threshold on presence rather than on sources. Given a 3.9% daily flip rate on presence and 92.4% churn on sources, alerting on source changes will page you almost every day and tell you nothing.

Presence is also only half the diagnosis. Whether anyone consistently wins a given prompt is a separate question, and on 33 of our 60 prompts nobody does.

Tracking tells you where you stand; it does not move you. Once the trend line shows you are absent on prompts you should own, the separate question of how to actually rank in ChatGPT is where to go next.

How many runs that takes is measurable rather than a matter of taste: on prompts that vary, one check misses the real rate by 33 points on average.

And give it a month before you read anything into the trend. With one in eight prompts genuinely unstable, a week of data is mostly noise around a small number of real movements.

Doing this manually versus not

A sixty-prompt set across four engines is 240 checks. At even fifteen seconds each that is an hour a day, every day, and the discipline required to do it identically each time is the part that fails first.

That is a reasonable thing to automate, and it is what GetIntel does. The honest framing is that automation buys you consistency and history rather than insight. The insight is in the one-in-eight, and you only find those by having run the same thing enough times to know what normal looks like.

Tags:AI visibilityChatGPTmonitoringresearch

Written by Tarang Agarwal

Tarang Agarwal is the founder of GetIntel. He writes about AI visibility, generative engine optimization, and growth for SaaS founders, marketing teams, and the agencies who run AI-search visibility as a service line.

FAQ

Frequently asked questions

In our tracking between 10 July and 10 August 2026, ChatGPT returned a different set of cited domains on 1,047 of 1,050 consecutive run pairs, or 99.7%. Runs were a median of 24.0 hours apart. Across all four engines measured the rate was 92.4% of 4,220 pairs, with Perplexity lowest at 76.4%.

Much less than the source churn implies. Across 4,360 consecutive run pairs measured between 10 July and 10 August 2026, our own brand flipped between cited and not-cited 169 times, a rate of 3.9%. By engine that ranges from 7.7% on ChatGPT down to 0.2% on Gemini. Sources are re-retrieved constantly while the conclusions drawn from them stay relatively stable.

For most prompts, yes. Of 520 prompt-and-engine series tracked between 10 July and 10 August 2026, 455 never changed state at all. The risk sits in the other 65, where a single check landed on the minority state 23.2% of the time. Since you cannot tell in advance which prompts are unstable, a schedule is what protects you.

Brand mentions. Across 4,220 comparable run pairs between 10 July and 10 August 2026, source sets changed on 92.4%, so alerting on them produces a notification almost every day with no signal in it. Presence flipped on 3.9% of the 4,360 consecutive pairs, which is rare enough that a change is worth looking at.

About a month. With roughly one prompt in eight genuinely unstable, a week of data mostly shows noise around a small number of real movements. Our own figures here come from 20 distinct run days spanning 10 July to 10 August 2026.

Put this into action

A Findability Score that refreshes daily, plus the exact fix, drafted and shipped through your coding agent. Built for founders, teams, and agencies.