Ask your vendor one question: does this number come from the API or from the interface a buyer actually uses? We sent the same buyer prompts down both paths on 11 August 2026 and compared what came back. A pair here means one prompt sent down both paths back to back. Across 11 clean pairs they shared 13.6% of their cited sources. The interesting part is what that number has to be compared against: the same ChatGPT interface, queried a day apart, shares only 16.7% of its sources with itself. Those two figures are close, and the honest reading is that from source overlap alone you cannot tell a different retrieval path from an ordinary different day. That is not something a dashboard will ever show you, because every tool renders the same tidy percentage regardless of where it came from.
What follows includes a result we got wrong first and corrected, because the correction is the most useful part. The full dataset is published, including the per-pair source lists, and nothing here was written into our tracked history. This was a read-only experiment.
Are the ChatGPT API and the interface the same thing?
OpenAI's Responses API with the web_search tool and the chatgpt.com interface are different products that happen to share a model. They retrieve separately, they surface citations differently, and neither is under any obligation to agree with the other.
Most AI visibility tools use the API, because it is cheap, fast and stable. Scraping the real interface costs more and breaks more often. That is a reasonable engineering trade, and it is usually invisible in the product you are sold.
It matters because your buyer is not using the API. They are typing into the interface. A score can be internally consistent, well defined and correctly calculated, and still be measuring a surface nobody buys from.
The result we got wrong
Our first batch was 5 pairs, of which 4 were clean. On those the interface averaged 3.50 cited sources against the API's 1.25, and every single clean pair pointed the same way. A 2.8x gap with unanimous direction looks like a finding.
It was not. We ran 7 more pairs and it vanished: 2.18 sources per answer from the API, 2.18 from the interface, across 11 pairs. The interface returned more in 4, the API returned more in 6, and they tied in 1.
We nearly published the first version. Four consistent pairs felt like enough because the direction was unanimous, and unanimity across a small sample is exactly what randomness looks like when you stop early. That is the same mistake as checking your brand in ChatGPT once and concluding you are invisible.
On source counts, then, this is a null result. Both paths return about the same number of sources.
So what actually differs between them?
The sources themselves. Across the 11 pairs, the API and the interface shared 13.6% of their cited domains by Jaccard overlap.

That number needs a comparison, and picking the wrong one is easy. Our churn dataset reports a pooled 23.6% overlap between consecutive runs, but that is an average across four engines and it is not a ChatGPT number. Broken out per engine, over 1,110 ChatGPT pairs:
ChatGPT source overlap compared across baselines, 10 July to 11 August 2026:
| comparison | shared sources |
|---|---|
| Perplexity vs itself, one day apart | 51.4% |
| Google AI Overviews vs itself | 16.9% |
| ChatGPT vs itself, one day apart | 16.7% |
| ChatGPT API vs ChatGPT interface, same moment | 13.6% |
| Gemini vs itself | 10.1% |
Perplexity's comparative stability, which shows up again in its citation-source patterns, is what drags the pooled average up to 23.6%, and quoting that as the ChatGPT baseline would have made our own result look far more dramatic than it is. Against the correct baseline the two figures are 16.7% and 13.6%, and with 11 pairs that difference is not one we would defend.
So the honest finding is a smaller and more useful one. Querying the API disagrees with the interface roughly as much as querying the interface again tomorrow does. Source overlap cannot separate the two, which means it is the wrong test to reach for, and any vendor claiming their API numbers match the interface cannot prove it this way either.
The disagreement that matters
Source lists are interesting. Whether you get named is the thing people actually buy a tool to find out, and on that the two paths disagreed too.
On 1 of the 11 pairs, the API cited us and the interface did not. That is a single case out of eleven and we are not going to build a rate on it. But it is an existence proof: the two paths can return different answers to the only question your dashboard is really reporting.
One disagreement in eleven is also enough to make the point that matters commercially. If a vendor reports your visibility from API data, some fraction of what they tell you is about a surface no buyer of yours will ever see.
Four questions to ask a vendor
None of these require technical knowledge, and all four have a right answer.
- Do you query the API or the rendered interface? If they use the API, that is not disqualifying, but it should be disclosed and it should be in the methodology page rather than something you have to ask for.
- How many runs is each number averaged over? A single run of anything is one draw from a distribution that changes almost every time. We measured a 99.7% source-change rate between consecutive ChatGPT runs.
- Can I see the raw answer text? If you cannot read the actual response that produced a score, you cannot audit the score. This is the fastest way to tell a measurement product from a dashboard.
- What happens when a run fails? Our own experiment made 24 calls across 12 pairs, and one of them came back empty on a deadline. We dropped that whole pair rather than scoring it. Counting it as zero citations would have quietly manufactured a gap between the two paths, and a tool that silently treats failures as absences will drift downward for reasons that have nothing to do with your brand.
What this does not establish
Eleven pairs, one engine, one day, one run each. This is a small experiment and it should be read as one.
The count result is a null finding rather than proof of equivalence: we did not detect a difference, which is not the same as showing there is none. With this sample only a large gap would have been visible, and the first batch shows how easily a small sample invents one.
The 13.6% overlap rests on firmer ground than the count result, because the baseline it is compared against comes from 1,110 ChatGPT pairs in a separate dataset. Even so, it is a single day and 11 pairs, and it is close enough to that baseline that we are describing an absence of separation rather than a measured gap.
We have also not tested Perplexity, Gemini or Google AI Overviews this way, and there is no reason to assume the gap is the same size on each. Our tracked set runs on the unattended daily path, and this experiment ran against the attended one.
How do I check my own tool?
If your tool exposes the raw answer for a prompt, take one and ask the same question in chatgpt.com by hand on the same day. You are not checking whether the wording matches, because it will not. You are checking whether the same sources appear.
If they mostly do, your tool is probably reading the surface your buyers read. If they mostly do not, you are looking at a different product's answer, and the number on your dashboard is measuring it rather than your market.
How do I just see what ChatGPT tells buyers, with no tool?
Open chatgpt.com and type the questions your buyers actually ask, in their words rather than yours. Not "best AI visibility tracker" but the sentence a person types, like "I run a small SaaS and I have no idea if ChatGPT recommends us, what should I use".
Run eight or ten of them. Read two things in each answer: which brands get named, and which sources the answer cites underneath. The brand list is the scoreboard; the sources are where the answer came from and therefore where the work is.
Then run the same set again a day later, because one pass tells you very little. We measured a 99.7% source-change rate between consecutive runs, so a source you saw once may simply not be there tomorrow. Anything that shows up in both passes is worth acting on; anything that appears once is a coin flip you happened to observe.
That costs an hour and needs no tool at all. It also will not show up as traffic: we earned 276 citations and 6 non-brand Google clicks in the same month. It is also exactly what a tracking product automates, which is the honest case for buying one rather than the exciting one.
GetIntel queries the rendered interface rather than the API on both its attended and unattended paths, which is the more expensive choice and the reason this comparison was cheap for us to run. That is a disclosure, not a recommendation. The four questions above are the useful part, and they apply to us as much as to anyone else.
