Guide

AI Citation Overlap: 1,283 Domains Across 5 Engines

We collected 491 AI answers and 1,283 cited domains across 5 AI engines. Only 1.1% of those domains appeared on all five. The full overlap matrix.

Tarang AgarwalAugust 2, 20269 min read
AI citation overlap study cover: 1,283 domains measured across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews

We collected 491 AI answers citing 1,283 distinct domains across ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews on 2 August 2026. Two-thirds of those domains appeared on exactly one engine. Just 14 — 1.1% — were cited by all five. Cross-engine overlap ranged from 4.8% to 15.7%, so winning on one engine predicts almost nothing about another.

The run, by the numbers:

  • 491 AI answers collected across 5 engines
  • 1,283 distinct domains cited in total
  • 856 (66.7%) of those domains appeared on exactly one engine
  • 14 (1.1%) domains were cited by all five engines
  • 4.8%–15.7% citation overlap between any two engines
  • 100 buyer-intent prompts spanning 10 verticals
  • Claude included — a measurement no competitor publishes at any scale

That number matters because most AI visibility tools, including ours, sell you a single cross-engine score. If the engines were broadly citing the same sources, one blended number would be a fair summary. They aren't, so it isn't — and we think the honest thing to do is publish the measurement rather than the marketing.

Disclosure: GetIntel is an AI visibility tracking product. We sell the thing this article complicates. Every number below comes from one run we did ourselves on a single day, and you can check our working — the full dataset is downloadable as CSV or JSON, and the method is at the end.

In this article

How much do AI engines actually cite the same sources?

Barely. Across every pair of engines we tested, citation overlap ran between 4.8% and 15.7% — meaning that even for the closest pair, more than four in five cited sources were unique to one engine or the other. The strongest agreement was Claude and Perplexity at 15.7%; the weakest were Claude/ChatGPT and Gemini/ChatGPT, both at 4.8%. No pair came close to half. We measured this as the mean per-prompt Jaccard score on cited domains: for each prompt, the set of domains engine A cited versus the set engine B cited, intersection over union, averaged across every prompt where both engines answered. Per-prompt matters — pooling across the whole corpus would inflate agreement by letting answers to unrelated questions match each other.

Engine pairPrompts comparedOverlapPrompts with zero shared domains
Claude ↔ Perplexity8115.7%11
Claude ↔ Gemini9213.2%34
Google AIO ↔ Perplexity7612.9%9
Gemini ↔ Perplexity7810.4%21
Gemini ↔ Google AIO859.6%29
Claude ↔ Google AIO888.6%30
Google AIO ↔ ChatGPT657.1%27
ChatGPT ↔ Perplexity605.8%24
Claude ↔ ChatGPT724.8%45
Gemini ↔ ChatGPT704.8%46
Overlap between engine pairs runs from 15.7% down to 4.8%. Every ChatGPT pair sits at the bottom.
Overlap between engine pairs runs from 15.7% down to 4.8%. Every ChatGPT pair sits at the bottom.

The right-hand column is the one that changed how we read this. These pairs don't merely agree weakly — on a large share of prompts they agree on nothing at all. Gemini and ChatGPT cited no domain in common on 46 of 70 comparable prompts. Claude and ChatGPT, 45 of 72. Same buyer question, same afternoon, entirely disjoint source sets.

Looking at it from the domain side rather than the pair side makes the same point more starkly:

Cited byDomainsShare
Exactly 1 engine85666.7%
Exactly 2 engines24218.9%
Exactly 3 engines1189.2%
Exactly 4 engines534.1%
All 5 engines141.1%
Two-thirds of the 1,283 cited domains appeared on just one engine. Only 14 appeared on all five.
Two-thirds of the 1,283 cited domains appeared on just one engine. Only 14 appeared on all five.

Fourteen domains made it onto every engine: amplitude.com, bitwarden.com, calendly.com, guideflow.com, gusto.com, learn.g2.com, infisical.com, monday.com, sequenzy.com, sparkhire.com, technologyadvice.com, techrepublic.com, thedigitalprojectmanager.com and zapier.com. That's the entire universally-cited set out of 1,283.

One number here is worth flagging honestly. Our ChatGPT ↔ Perplexity figure of 5.8% is identical to the figure Wellows published from a far larger sample of 571,729 answers. We did not expect to land on their number to the decimal, and we can't claim a clean replication because they don't publish their exact overlap formula — a like-for-like comparison isn't verifiable from what's public. But it held at two different sample sizes in our own run (n=55 before a retry pass, n=60 after), so it isn't an artifact of a thin sample on our side.

Why does ChatGPT overlap least with everything else?

Because it draws from a far smaller pool of sources than the others. ChatGPT cited 205 distinct domains across our 100 prompts, against 630 for Google AI Overviews — roughly a third as many. It isn't only citing fewer sources per answer; it returns to the same narrow set repeatedly, across different questions and different verticals. Any two engines drawing from small, differently-shaped pools will overlap less by construction, and ChatGPT has the smallest pool in the set. The full spread:

EngineAnsweredDomains per answerDistinct domains across 100 prompts
Google AI Overviews91/10010.05630
Perplexity100/1007.61457
Claude100/1005.27405
ChatGPT100/1003.66205
Gemini100/1003.50279

ChatGPT cites from a pool roughly a third the size of AI Overviews'. It also, we found separately, names more brands than any other engine while citing the fewest sources — recommending confidently and linking sparingly. It isn't just citing fewer sources per answer — it returns to the same narrow set of sources repeatedly across different questions and different verticals. Gemini cites slightly fewer domains per answer but draws them from a wider pool.

The practical consequence: ChatGPT has the fewest available slots. Being the eleventh-best source for a topic can still earn a Google AI Overviews citation, because AIO is citing ten domains an answer from a pool of 630. The same position earns nothing on ChatGPT. That makes ChatGPT visibility a materially harder problem than the others, not the same problem on a different surface.

For what it's worth, our Perplexity figure of 7.61 domains per answer closely matches the 7.5 that Omnia reported from 42 million citations, while contradicting the 4.82 that Wellows reported. We're adding a third data point to a contested number rather than settling it.

What does Claude cite that the other engines don't?

Specialist trade publications and vendor documentation — and, uniquely among the five, almost no aggregators. The most-repeated claim in this category is that AI search engines cite Reddit, YouTube and LinkedIn above all else. For Google AI Overviews our data supports that emphatically: it cited YouTube on 56 of the 91 prompts it answered and Reddit on 54. For Claude it is simply false. Not one aggregator appears anywhere in Claude's top sources.

Here is the most-cited domain for each engine, by how many distinct prompts it appeared on:

EngineTop sources
Google AI Overviewsyoutube.com (56 of 91 prompts), reddit.com (54), zapier.com (15)
Perplexityzapier.com (25), reddit.com (16), techradar.com (12)
ChatGPTreddit.com (16), techradar.com (15), hubspot.com (9)
Geminipcmag.com (7), learn.g2.com (5), gartner.com (5)
Claudeamplitude.com (7), thedigitalprojectmanager.com (5), ventureharbour.com (5)

Claude is the only engine in the set with no aggregator anywhere in its top sources. Its citations skew toward vendor documentation and specialist trade publications, distributed across a long tail — 405 distinct domains, none appearing on more than 7 of 100 prompts. Gemini leans on traditional technology media. Google AI Overviews cited YouTube or Reddit on well over half of all prompts it answered.

This is why "get mentioned on Reddit" is engine-specific advice being sold as universal advice. It is a reasonable strategy for AI Overviews. It is close to irrelevant for Claude. If you want a fuller treatment of Claude specifically, we've written separately about how to rank in Claude.

Does a single cross-engine visibility score mean anything?

We've argued before that averaging engines together hides the real problem. This run is the evidence for that argument rather than a restatement of it.

A blended score is an average over five largely independent measurements. If you score 40, that could mean a consistent 40 everywhere, or 100 on Perplexity and near-zero on the other four. Those are completely different situations calling for completely different work, and one number cannot distinguish them.

That doesn't make a composite useless — it makes it a headline that requires a breakdown underneath. A single score is defensible as a trend line you watch move over time. It is not defensible as a diagnosis. Any tool that shows you the composite without letting you open up the per-engine detail, ours included, is hiding the part you actually need to act on.

What do people get wrong about this?

Treating one engine's result as a proxy for the rest. The most common version is checking ChatGPT, seeing a brand mentioned, and assuming coverage. On our data ChatGPT is the worst possible proxy — it has the lowest overlap with every other engine in the set.

Assuming a citation win transfers. Two-thirds of cited domains appeared on exactly one engine. The base rate for transfer is low, so earning a citation should be treated as engine-specific until measured otherwise.

Applying source-mix advice universally. Reddit and YouTube strategy is close to optimal for AI Overviews and close to useless for Claude. Advice with no engine attached is advice averaged across engines that behave differently.

Reading a one-off snapshot as a constant. This includes ours. Profound's data shows 40-60% of cited domains change month-to-month for identical questions, so a single run is a photograph, not a law. Anything you conclude from one measurement should be re-measured before you build a quarter's work on it.

Comparing overlap figures without checking the metric. Published overlap numbers in this category range from 5.8% to 12% partly because vendors compute them differently and mostly don't publish the formula. Ours is defined above precisely so it can be argued with.

How did we measure this?

Date: 2 August 2026. Sample: 100 buyer-intent prompts across 10 verticals — customer support, CRM, project management, email marketing, HR and payroll, accounting, ecommerce, security, analytics, and meetings and scheduling — 10 prompts each, all unique. Answers collected: 491.

A verbatim example prompt: "What is the best customer support software for a small SaaS team in 2026?"

Engines and surfaces: ChatGPT, Perplexity and Gemini were captured by scraping the real consumer interfaces. Google AI Overviews came through a SERP data provider. Claude was probed through the Anthropic API's web search tool, because no equivalent consumer-interface capture exists for Claude. That asymmetry is real and we'd rather state it than bury it: Claude's numbers here describe the API surface, and a consumer using claude.ai may see something different. Claude ran on Sonnet 5.

Two limitations worth stating plainly. Our sample is 100 prompts; Wellows' comparable work covers 35,404 queries. Ours is small enough that individual pair figures carry meaningful uncertainty, and we'd treat the overall pattern as much more reliable than any single cell. Second, Google AI Overviews answered 91 of 100 prompts — the nine misses are queries where Google rendered no AI Overview at all, which is a real result rather than a failure, and they're excluded from both the numerator and denominator rather than counted as zeros.

The data. Every figure in this article can be recomputed from the dataset we've published: ai-citation-overlap-2026-08-02.csv (320 KB) or the same data as JSON with the run metadata attached. It's one row per prompt, engine and cited domain — 2,980 rows covering all 491 answers. We've deliberately published the citation layer rather than the engines' verbatim answer text: it's what every number here rests on, so it's sufficient to check us, without redistributing several hundred thousand words of other companies' model output. Free to use with attribution.

We intend to re-run this quarterly, because a citation snapshot ages quickly and a number without a date is how this category ended up with five irreconcilable figures for the same measurement.

What should you do with this data?

Stop asking "are we visible in AI search". It's the wrong unit. Ask which engines you're visible in, which you aren't, and whether the sources being cited in your category on that specific engine are ones you could realistically appear on.

That's the job GetIntel does — tracking your prompts across engines separately and showing which sources each one is actually pulling from, rather than collapsing it into one number. If you want to see the per-engine picture for your own brand, start with a free report, or read our honest roundup of AI visibility tools — including the ones we lose to — if you're still deciding what to use.

Tags:AI visibilityGEOClaudeChatGPTresearch

Written by Tarang Agarwal

Tarang Agarwal is the founder of GetIntel. He writes about AI visibility, generative engine optimization, and growth for SaaS founders, marketing teams, and the agencies who run AI-search visibility as a service line.

FAQ

Frequently asked questions

In our 2 August 2026 run of 100 buyer prompts across five engines, pairwise cited-domain overlap ranged from 4.8% to 15.7%. Two-thirds of the 1,283 distinct domains cited appeared on exactly one engine, and only 14 domains (1.1%) were cited by all five.

ChatGPT, on this data. It cited 205 distinct domains across 100 prompts against Google AI Overviews' 630, meaning it returns to a much narrower set of sources. Fewer available slots makes it a harder target, and it also had the lowest citation overlap with every other engine.

It depends entirely on the engine. Google AI Overviews cited YouTube on 56 of 91 prompts and Reddit on 54. Claude cited no aggregator at all in its top sources, favouring vendor documentation and specialist trade publications instead. Universal source-mix advice averages over engines that behave very differently.

As a trend line, yes. As a diagnosis, no. A blended score averages five largely independent measurements, so a score of 40 could mean a consistent 40 everywhere or a near-perfect result on one engine and near-zero on four others. Those need different work, and one number cannot tell them apart.

100 unique buyer-intent prompts across 10 verticals, run on 2 August 2026, producing 491 answers. Overlap is the mean per-prompt Jaccard score on cited domains, averaged across prompts where both engines answered. ChatGPT, Perplexity and Gemini were captured from their consumer interfaces; Google AI Overviews via a SERP provider; Claude via the Anthropic API's web search tool, as no consumer-interface capture exists for Claude.

Put this into action

A Findability Score that refreshes daily, plus the exact fix, drafted and shipped through your coding agent. Built for founders, teams, and agencies.