Guide

Does Your AI Visibility Score Reflect the Real Interface?

We sent the same buyer prompts through ChatGPT's API and the real chatgpt.com interface. They shared 13.6% of their sources, less than the interface shares with itself.

Tarang AgarwalAugust 11, 20267 min read min read
Title card for an article on whether AI visibility scores reflect the real ChatGPT interface or the API

Ask your vendor one question: does this number come from the API or from the interface a buyer actually uses? We sent the same buyer prompts down both paths on 11 August 2026 and compared what came back. A pair here means one prompt sent down both paths back to back. Across 11 clean pairs they shared 13.6% of their cited sources. The interesting part is what that number has to be compared against: the same ChatGPT interface, queried a day apart, shares only 16.7% of its sources with itself. Those two figures are close, and the honest reading is that from source overlap alone you cannot tell a different retrieval path from an ordinary different day. That is not something a dashboard will ever show you, because every tool renders the same tidy percentage regardless of where it came from.

What follows includes a result we got wrong first and corrected, because the correction is the most useful part. The full dataset is published, including the per-pair source lists, and nothing here was written into our tracked history. This was a read-only experiment.

Are the ChatGPT API and the interface the same thing?

OpenAI's Responses API with the web_search tool and the chatgpt.com interface are different products that happen to share a model. They retrieve separately, they surface citations differently, and neither is under any obligation to agree with the other.

Most AI visibility tools use the API, because it is cheap, fast and stable. Scraping the real interface costs more and breaks more often. That is a reasonable engineering trade, and it is usually invisible in the product you are sold.

It matters because your buyer is not using the API. They are typing into the interface. A score can be internally consistent, well defined and correctly calculated, and still be measuring a surface nobody buys from.

The result we got wrong

Our first batch was 5 pairs, of which 4 were clean. On those the interface averaged 3.50 cited sources against the API's 1.25, and every single clean pair pointed the same way. A 2.8x gap with unanimous direction looks like a finding.

It was not. We ran 7 more pairs and it vanished: 2.18 sources per answer from the API, 2.18 from the interface, across 11 pairs. The interface returned more in 4, the API returned more in 6, and they tied in 1.

We nearly published the first version. Four consistent pairs felt like enough because the direction was unanimous, and unanimity across a small sample is exactly what randomness looks like when you stop early. That is the same mistake as checking your brand in ChatGPT once and concluding you are invisible.

On source counts, then, this is a null result. Both paths return about the same number of sources.

So what actually differs between them?

The sources themselves. Across the 11 pairs, the API and the interface shared 13.6% of their cited domains by Jaccard overlap.

Bar chart comparing source overlap. Perplexity shares 51.4% of its sources with itself a day apart, Google AI Overviews 16.9%, ChatGPT 16.7% and Gemini 10.1%, against 13.6% between ChatGPT's API and its interface.
Bar chart comparing source overlap. Perplexity shares 51.4% of its sources with itself a day apart, Google AI Overviews 16.9%, ChatGPT 16.7% and Gemini 10.1%, against 13.6% between ChatGPT's API and its interface.

That number needs a comparison, and picking the wrong one is easy. Our churn dataset reports a pooled 23.6% overlap between consecutive runs, but that is an average across four engines and it is not a ChatGPT number. Broken out per engine, over 1,110 ChatGPT pairs:

ChatGPT source overlap compared across baselines, 10 July to 11 August 2026:

comparisonshared sources
Perplexity vs itself, one day apart51.4%
Google AI Overviews vs itself16.9%
ChatGPT vs itself, one day apart16.7%
ChatGPT API vs ChatGPT interface, same moment13.6%
Gemini vs itself10.1%

Perplexity's comparative stability, which shows up again in its citation-source patterns, is what drags the pooled average up to 23.6%, and quoting that as the ChatGPT baseline would have made our own result look far more dramatic than it is. Against the correct baseline the two figures are 16.7% and 13.6%, and with 11 pairs that difference is not one we would defend.

So the honest finding is a smaller and more useful one. Querying the API disagrees with the interface roughly as much as querying the interface again tomorrow does. Source overlap cannot separate the two, which means it is the wrong test to reach for, and any vendor claiming their API numbers match the interface cannot prove it this way either.

The disagreement that matters

Source lists are interesting. Whether you get named is the thing people actually buy a tool to find out, and on that the two paths disagreed too.

On 1 of the 11 pairs, the API cited us and the interface did not. That is a single case out of eleven and we are not going to build a rate on it. But it is an existence proof: the two paths can return different answers to the only question your dashboard is really reporting.

One disagreement in eleven is also enough to make the point that matters commercially. If a vendor reports your visibility from API data, some fraction of what they tell you is about a surface no buyer of yours will ever see.

Four questions to ask a vendor

None of these require technical knowledge, and all four have a right answer.

  • Do you query the API or the rendered interface? If they use the API, that is not disqualifying, but it should be disclosed and it should be in the methodology page rather than something you have to ask for.
  • How many runs is each number averaged over? A single run of anything is one draw from a distribution that changes almost every time. We measured a 99.7% source-change rate between consecutive ChatGPT runs.
  • Can I see the raw answer text? If you cannot read the actual response that produced a score, you cannot audit the score. This is the fastest way to tell a measurement product from a dashboard.
  • What happens when a run fails? Our own experiment made 24 calls across 12 pairs, and one of them came back empty on a deadline. We dropped that whole pair rather than scoring it. Counting it as zero citations would have quietly manufactured a gap between the two paths, and a tool that silently treats failures as absences will drift downward for reasons that have nothing to do with your brand.

What this does not establish

Eleven pairs, one engine, one day, one run each. This is a small experiment and it should be read as one.

The count result is a null finding rather than proof of equivalence: we did not detect a difference, which is not the same as showing there is none. With this sample only a large gap would have been visible, and the first batch shows how easily a small sample invents one.

The 13.6% overlap rests on firmer ground than the count result, because the baseline it is compared against comes from 1,110 ChatGPT pairs in a separate dataset. Even so, it is a single day and 11 pairs, and it is close enough to that baseline that we are describing an absence of separation rather than a measured gap.

We have also not tested Perplexity, Gemini or Google AI Overviews this way, and there is no reason to assume the gap is the same size on each. Our tracked set runs on the unattended daily path, and this experiment ran against the attended one.

How do I check my own tool?

If your tool exposes the raw answer for a prompt, take one and ask the same question in chatgpt.com by hand on the same day. You are not checking whether the wording matches, because it will not. You are checking whether the same sources appear.

If they mostly do, your tool is probably reading the surface your buyers read. If they mostly do not, you are looking at a different product's answer, and the number on your dashboard is measuring it rather than your market.

How do I just see what ChatGPT tells buyers, with no tool?

Open chatgpt.com and type the questions your buyers actually ask, in their words rather than yours. Not "best AI visibility tracker" but the sentence a person types, like "I run a small SaaS and I have no idea if ChatGPT recommends us, what should I use".

Run eight or ten of them. Read two things in each answer: which brands get named, and which sources the answer cites underneath. The brand list is the scoreboard; the sources are where the answer came from and therefore where the work is.

Then run the same set again a day later, because one pass tells you very little. We measured a 99.7% source-change rate between consecutive runs, so a source you saw once may simply not be there tomorrow. Anything that shows up in both passes is worth acting on; anything that appears once is a coin flip you happened to observe.

That costs an hour and needs no tool at all. It also will not show up as traffic: we earned 276 citations and 6 non-brand Google clicks in the same month. It is also exactly what a tracking product automates, which is the honest case for buying one rather than the exciting one.

GetIntel queries the rendered interface rather than the API on both its attended and unattended paths, which is the more expensive choice and the reason this comparison was cheap for us to run. That is a disclosure, not a recommendation. The four questions above are the useful part, and they apply to us as much as to anyone else.

Tags:AI visibilityChatGPTmethodologyresearch

Written by Tarang Agarwal

Tarang Agarwal is the founder of GetIntel. He writes about AI visibility, generative engine optimization, and growth for SaaS founders, marketing teams, and the agencies who run AI-search visibility as a service line.

FAQ

Frequently asked questions

No. On 11 August 2026 we sent the same buyer prompts through OpenAI's Responses API with web_search and through the real chatgpt.com interface. Across 11 clean pairs they shared only 13.6% of their cited domains. The right comparison is ChatGPT against itself a day apart, which is 16.7% over 1,110 pairs. Those are close, so source overlap alone cannot tell a different path from a different day.

No, on our evidence. Across 11 pairs measured on 11 August 2026 both averaged 2.18 cited sources per answer. Our first batch of 4 clean pairs suggested a 2.8x gap in the interface's favour, but that disappeared entirely when we ran 7 more pairs. It was noise from a small sample.

Yes. On 1 of 11 pairs measured on 11 August 2026, the API cited our domain and the interface did not. One case in eleven is too few to build a rate on, but it demonstrates the two paths can return different answers to the exact question a visibility dashboard reports.

Ask directly, and ask to see the raw answer text behind a score. If you can read the response, take one prompt and run it by hand in chatgpt.com the same day, then compare which sources appear rather than whether the wording matches. Wording will differ either way; overlapping sources are the signal.

More than one, because a single run is one draw from a distribution that moves constantly. We measured a 99.7% source-change rate between consecutive ChatGPT runs a median 24 hours apart, across 1,050 ChatGPT pairs between 10 July and 10 August 2026.

Type eight to ten real buyer questions into chatgpt.com in the words a buyer would use, and record which brands are named and which sources are cited. Then repeat the set a day later. We measured a 99.7% source-change rate between consecutive ChatGPT runs across 1,050 ChatGPT pairs between 10 July and 10 August 2026, so only what appears in both passes is worth acting on.

Put this into action

A Findability Score that refreshes daily, plus the exact fix, drafted and shipped through your coding agent. Built for founders, teams, and agencies.