Across the ten software categories we tested, there is no top-10 list of AI sources to get into. We put 100 buying questions to five AI engines on 2 August 2026 and logged all 3,754 citations they returned. Those citations went to 1,258 different domains. The single most-cited domain in the entire set accounted for 3.5% of them, and the top ten together accounted for under 15%.
That matters because the standard advice in this category is to identify the handful of sources AI engines trust and get yourself into them. Our data says that handful does not exist — at least not in the ten mainstream software categories we tested.
Disclosure: GetIntel sells AI visibility tracking, so we have an interest in how this gets measured. Every figure below comes from one run we did ourselves, and the full per-domain table is downloadable as CSV and JSON — all 1,258 domains, not a top-20 summary — so you can recompute any of it.
In this article
- Which domains do AI engines cite most?
- Is there a short list of sources worth targeting?
- Does concentration differ by engine?
- So which is the number one source, YouTube or Reddit?
- What do people get wrong about this?
- How did we measure this?
- What should you do instead?
Which domains do AI engines cite most?
YouTube, Reddit and Zapier — but none of them by much. YouTube is the most-cited domain in the set at 132 citations, which is 3.5% of the 3,754 we logged. Reddit follows at 3.0% and Zapier at 1.7%. After that it drops below 1% almost immediately. These are the ten most-cited domains across all five engines:
| Domain | Citations | Share of all citations | Answers it appeared in |
|---|---|---|---|
| youtube.com | 132 | 3.5% | 68 of 491 |
| reddit.com | 114 | 3.0% | 88 of 491 |
| zapier.com | 64 | 1.7% | 50 of 491 |
| g2.com | 43 | 1.1% | 31 of 491 |
| pcmag.com | 37 | 1.0% | 25 of 491 |
| techradar.com | 35 | 0.9% | 29 of 491 |
| sequenzy.com | 34 | 0.9% | 18 of 491 |
| gusto.com | 32 | 0.9% | 19 of 491 |
| shopify.com | 32 | 0.9% | 21 of 491 |
| amplitude.com | 32 | 0.9% | 23 of 491 |
Two things worth noticing. The list is not what the category talks about — it mixes an aggregator, a forum, a product blog, two review sites, two trade publications and three vendors' own sites. And seventh place is sequenzy.com, a domain almost nobody in this space would name if asked to list the sources AI trusts.
Is there a short list of sources worth targeting?
Not in these ten categories. The top ten domains account for 14.8% of all citations, and you have to go to roughly 300 domains before you cover two-thirds of them.

| Top N domains | Share of all citations |
|---|---|
| Top 1 | 3.5% |
| Top 3 | 8.3% |
| Top 10 | 14.8% |
| Top 25 | 22.8% |
| Top 50 | 31.1% |
| Top 100 | 43.0% |
| Top 300 | 66.9% |
The tail is the story. 721 of the 1,258 domains — 57% — were cited exactly once across the whole run. Those single-citation domains still carry about a fifth of all citations between them. This is not a distribution with a head you can capture; it is a distribution where being cited at all is a wide, shallow lottery, and where most of the winners appear once.
That does not mean sources are irrelevant. It means "get into the sources AI trusts" is not a plan, because the set is too large and too flat to enumerate, and because most of it turns over — a point our comparison of what different engines cite makes from the other direction.
Does concentration differ by engine?
Sharply — by a factor of 5.4. Google AI Overviews puts 19.6% of its citations into just three domains, while Claude's top three account for 3.6%. So "is there a short list?" has a different answer depending on which engine you ask about.

| Engine | Top 3 domains' share | Most-cited domain | Distinct domains cited |
|---|---|---|---|
| Google AI Overviews | 19.6% | youtube.com (11.0%) | 621 |
| ChatGPT | 12.4% | techradar.com (5.2%) | 198 |
| Perplexity | 7.3% | zapier.com (3.3%) | 453 |
| Gemini | 5.7% | ventureharbour.com (2.0%) | 278 |
| Claude | 3.6% | amplitude.com (1.4%) | 403 |
Google AI Overviews is the one engine where source targeting is a coherent strategy, and it is coherent mainly because it leans so heavily on YouTube — 11% of everything it cites. Claude is the opposite: no domain reaches 1.5%, and its most-cited source is a product site rather than a publisher.
ChatGPT is the interesting middle case. It cites the fewest distinct domains of any engine (198) yet its top three only reach 12.4%, because a quarter of its answers cite nothing at all — a pattern we measured separately when we counted how many sources each engine cites per answer.
So which is the number one source, YouTube or Reddit?
Both, depending on which denominator you use — and that ambiguity is why a single "top source" figure should always be treated with suspicion.
YouTube wins on citation count: 132 against Reddit's 114. Reddit wins on reach: it appeared in 88 of 491 answers where YouTube appeared in 68. YouTube gets cited repeatedly inside the same answer; Reddit gets cited once across more answers.
Neither ranking is wrong. They answer different questions — "which domain do engines cite most often" versus "which domain is a buyer most likely to encounter". A vendor publishing "Reddit is the number one AI source" and one publishing "YouTube is" can both be honest, and neither usually says which denominator it used. We wrote about that failure mode across the category yesterday; this is a clean example of it inside our own data.
What do people get wrong about this?
Treating source targeting as the whole strategy, assuming the well-known sources dominate, and reading one engine's behaviour as the general case. The first is the expensive one.
Treating source targeting as the whole strategy
If the top ten domains carry under 15% of citations, then winning a place in all ten still leaves 85% of the citation surface untouched. Source targeting is worth doing where it is cheap — a G2 profile, a Reddit presence — but it cannot be the plan.
Assuming the well-known sources dominate
They do not. sequenzy.com outranks most names the category discusses, and Claude's most-cited domain is a product site. Reputation inside the industry is a poor guide to what engines actually pull.
Reading one engine's behaviour as the general case
Google AI Overviews is five times more concentrated than Claude. Any advice derived from watching a single engine — including "AI loves YouTube", which is largely a statement about AI Overviews — generalises badly.
How did we measure this?
Date: 2 August 2026. Sample: 100 buyer-intent questions across 10 software categories, 10 each, put to five engines. Answers: 491. Citations: 3,754 across 1,258 distinct domains.
Counting. Domains were collapsed to a registrable form, so learn.g2.com and g2.com count as one publisher. That choice matters and it works against our own thesis: not collapsing them would inflate the domain count and make the tail look even longer than it is.
Two rankings. We report both citations and answer-reach for every domain because they disagree at the top, and reporting only one would hide that.
Engines and surfaces. ChatGPT, Perplexity and Gemini were captured from their consumer interfaces; Google AI Overviews through a SERP data provider; Claude through the Anthropic API's web search tool, because no consumer-interface capture exists for it. Claude's figures therefore describe the API surface, and Claude ran on Sonnet 5.
Limitations. Ten software categories are not the whole internet, and concentration is likely to differ in categories with a dominant reference site — this is one reason we would not quote these percentages for, say, healthcare or law. Google AI Overviews answered 91 of 100 questions; the nine misses are excluded rather than counted as zero. This is one day's snapshot of systems that change monthly.
Checking us. ai-source-concentration-2026-08-02.csv lists all 1,258 domains with citations, share, answer-reach and how many engines cited each, plus a JSON summary. The per-engine table has its own file — every domain each engine cited — because the aggregate table alone could not reproduce it, and a promise that you can recompute our figures should be true of all of them. Re-runs will ship under their own date rather than overwriting this one.
What should you do instead?
Stop trying to get into a list, and start finding the questions where you are already close.
The practical version: take the ten or twenty questions your buyers actually ask, look at what gets cited on each one, and work on the specific pages that already rank near the citation boundary. That is a smaller and more tractable job than earning placement on a canonical source list, and unlike the list, it exists.
Three things follow from the data above. On Google AI Overviews, video is worth taking seriously — 11% of its citations go to YouTube, which is a real concentration and a real opportunity. On Claude and Gemini, no source-targeting strategy will move much, so the work is being the thing that gets named rather than linked. And on any engine, checking which sources actually appear for your own category beats generalising from someone else's.
If you would rather not assemble that by hand, that is what GetIntel does — it tracks the questions your buyers ask across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews and shows which sources each engine pulled, per question, with the date attached.
