Ask an AI engine which tool to buy in a category, and roughly half its answer is built from the tool vendors' own blogs. Across the 40 most-cited domains in the AI-visibility category, 9,049 citations from 3,244 answers, 51.7% of citations point at companies selling a product in that category. Independent journalism accounts for 5.7% (GetIntel source report, 6 August 2026).
Disclosure: GetIntel sells an AI visibility tool, so we are one of the vendors counted here, 1.7% of those citations are ours. The per-domain classification is published as CSV and JSON so you can disagree with how we labelled any of it and recompute.
In this article
- Who actually writes the sources AI cites?
- Which individual pages are doing the work?
- Is this vendors gaming the system?
- What does this mean if you are buying?
- Does our own tool undercount vendor influence?
- How did we measure it?
Who actually writes the sources AI cites?
Mostly the vendors themselves. Here is the full breakdown of who wrote the pages cited across those 3,244 answers:
| who wrote the source | citations | share |
|---|---|---|
| Companies selling a tool in this category | 4,675 | 51.7% |
| Reddit, YouTube, LinkedIn, Medium, Instagram | 2,350 | 26.0% |
| Agency and consultancy content | 563 | 6.2% |
| Other corporate sites (Zapier, HubSpot, GitHub) | 622 | 6.9% |
| Independent editorial (TechRadar, Search Engine Land) | 514 | 5.7% |
| Academic papers (arXiv) | 325 | 3.6% |
25 of the 40 most-cited domains are companies selling into the category they are being cited about.

Split further, 41.9% goes to dedicated AI-visibility vendors and 9.7% to SEO suites with an AI module. Whether Semrush counts as a "vendor" here is arguable, which is why the two are separated and the per-domain labelling is published rather than asserted.
Which individual pages are doing the work?
A handful of listicles, most of them written by companies that appear in their own list. Across the same AI-visibility dataset, 9,049 citations from 3,244 AI answers on 6 August 2026, ranked by citations to a single URL: zapier.com/blog/best-ai-visibility-tool at 154, trustmary.com/ai-visibility/best-ai-search-visibility-tools at 150, techradar.com's tracking guide at 144, seranking.com/ai-visibility-tracker.html at 120 and semrush.com/blog/best-ai-visibility-tools at 114.

Two things stand out. The concentration is extreme, a single page can carry more citations than most domains do in total. And the two highest are corporate blogs from companies that do not sell an AI-visibility tool at all, which is its own signal: in a category this young, general-purpose domain authority beats subject-matter specificity.
Is this vendors gaming the system?
No, and that is the more uncomfortable answer. Nothing here shows manipulation. It shows that in a category two years old, vendor marketing is most of what has been written.
Someone had to publish "the best AI visibility tools" before an engine could cite anything. The people with the strongest incentive to write that were the people selling the tools. Independent outlets have covered the space, but thinly, TechRadar and Search Engine Land together account for 5.7% of citations, and much of that sits in a handful of pages.
Engines are not choosing vendor content over independent content. They are choosing the content that exists.
One detail makes the thinness concrete: YouTube carries 758 of the 9,049 citations, and 737 of those are a single video. Not a broad video presence, one upload, doing almost all of the work. A source layer that thin is easy to become part of, and easy to mistake for consensus.
What does this mean if you are buying?
Treat an AI tool recommendation as a summary of vendor marketing, not as a verdict. We took this a step further the next day and checked the roundups themselves: 16 of 17 vendor-written "best tools" lists rank their own product first. Three practical consequences:
-
The tools named most are the ones that published most. Citation frequency in a young category measures publishing volume and domain age far more than product quality.
-
Ask where a claim came from. If an engine tells you tool X is best for Y, the underlying source is more likely to be X's own blog, or a listicle by a competitor of X, than an independent test.
-
Weight the 26% differently. Reddit, YouTube and LinkedIn are the largest non-vendor block, and Reddit's most-cited threads here are people arguing about which tools are overpriced. That is closer to independent signal than anything in the 51.7%.
-
Check whether the recommendation survives the question being rephrased. Vendor listicles are written against a handful of high-volume phrasings. Ask the same buying question three different ways, by job to be done, by budget, by the engine you care about, and the names that persist across all three are less likely to be an artefact of who optimised for one phrase.
A concrete way to test this yourself: take the tool an engine recommended, ask it where that recommendation came from, and look at whether the sources it names sell a competing product. In this category, on these numbers, better than even odds say at least one does.
If you want a genuinely independent read, the uncomfortable answer is that it mostly does not exist yet in this category, and no tool, ours included, can conjure it. What exists instead is the raw material: the sources each engine actually pulled, which at least lets you see whose blog is behind an answer before you act on it.
Does our own tool undercount vendor influence?
GetIntel's source report classifies citations as yours, competitor, or third-party, and for our own account it reads 1% yours, 8% competitor, 91% third-party. Those three come from our own account dashboard rather than the dataset above, so unlike every other figure here you cannot recompute them from the CSV. That 91% looks like a healthy independent majority.
It is not. "Competitor" only counts the competitors an account has explicitly tracked. Every other vendor (and there are dozens) falls into "third party" alongside Reddit and arXiv. So our own product reports 8% competitor influence in a category where the real vendor share is above 50%.
That is a real limitation in something we sell, found by running our own analysis against our own dashboard. The number is not wrong for what it measures; it is just measuring tracked competitors rather than commercial interest, and those are very different things.
How did we measure it?
Source. GetIntel's live source report on 6 August 2026, covering 3,244 AI answers across ChatGPT, Perplexity, Gemini and Google AI Overviews, on buyer-intent questions in the AI-visibility category. Every citation figure here is summed from the per-domain rows in the published CSV, which total 9,049. The report's own header reads 9,047 total sources: a two-citation gap between the aggregate and the sum of its parts that we could not reconcile, and are leaving visible rather than rounding away.
Scope. The 40 most-cited domains, which between them carry 9,049 citations. The long tail below that is not classified.
Classification. Each domain hand-labelled as vendor, SEO suite, UGC, agency content, editorial, academic or other corporate. This is the judgement call in the analysis, so the per-domain labels ship as CSV. Reasonable people will move Semrush, Ahrefs and Frase between buckets; the vendor share stays above half in any labelling we tried.
A denominator warning, including on our own dashboard. Every percentage here is citations divided by the 9,049 citations in the sample. That is not the only denominator available: GetIntel's own source report shows Reddit at "27%", which is its 889 citations divided by the 3,244 answers. That is a citations-per-answer rate, not a share of citations. Reddit's actual share of citations is 9.8%. The two numbers differ by a factor of nearly three, and we caught ourselves about to publish the wrong one.
Limitations. One category, ours, which is unusually vendor-dense because the buyers are marketers. A consumer category would likely look different. Citation counts are not weighted by how prominently a source was used in an answer. And this is a rolling window that moves: the figures above are 6 August and will drift.
What this changes for us
We are 1.7% of the problem we just described, and the honest position is that publishing more will make us a larger part of it rather than a smaller one.
What we can control is whether our contribution is worth citing. Our own five-engine citation study is built the same way. Every research piece we publish ships the raw data, dates its figures, and names the limitation, because the thing this category is short of is not more vendor content, it is anything a reader can check. You can start free if you want to run this on your own category, and our AI visibility tracking shows which sources each engine pulled for you specifically.
If you are on the receiving end of this, we cover why ChatGPT cites competitors instead of your brand.
