We ran 100 buying questions through ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews on 2 August 2026 and counted two things separately: how many real vendors each engine named, and how many sources it cited. The two barely relate. ChatGPT named the most brands while citing the fewest sources. Google AI Overviews did the exact opposite.
Most AI visibility tools measure one or the other and report it as "visibility". Our data says that's a category error — naming and citing are different behaviours, and the ratio between them swings roughly fourfold depending on which engine you ask.
Disclosure: GetIntel sells AI visibility tracking, so we have an interest here. Every number below comes from one run we did ourselves, and both halves of the data are downloadable — which brands each engine named and which sources each engine cited — so you can check the arithmetic yourself.
In this article
- What did we actually run?
- How many brands does each engine name?
- Does this match what's already been published?
- Which engine is most likely to name nobody at all?
- What does this change about measuring AI visibility?
- What do people get wrong about this?
- How did we measure this?
- What should you do with this?
What did we actually run?
100 buyer-intent questions — the kind someone types when they're about to spend money, like "What is the best CRM for a small B2B sales team?" — spread across 10 software categories: customer support, CRM, project management, email marketing, HR and payroll, accounting, ecommerce, security, analytics, and meetings and scheduling. Ten questions each. Every question went to all five engines on the same day, producing 491 answers.
Then we counted brand names in the answer text against a hand-built list of 223 real vendors, and counted the domains each engine cited, per answer.
How many brands does each engine name?
Between 3.59 and 5.28 per answer, and the ranking is nearly the reverse of how many sources they cite.
| Engine | Brands named per answer | Sources cited per answer | Named per source cited |
|---|---|---|---|
| ChatGPT | 5.28 | 3.66 | 1.44 |
| Gemini | 4.53 | 3.50 | 1.29 |
| Claude | 4.22 | 5.27 | 0.80 |
| Perplexity | 3.78 | 7.61 | 0.50 |
| Google AI Overviews | 3.59 | 10.05 | 0.36 |

ChatGPT names 1.44 brands for every source it links. Google AI Overviews names 0.36. That's a fourfold spread in the same measurement, on identical questions asked the same day.
ChatGPT behaves like someone answering from memory: it recommends confidently and links sparingly. AI Overviews behaves like a bibliography — lots of sources, comparatively few firm recommendations. Claude sits almost exactly in balance at 0.80.
This matters because it decides what "share of voice" even means. On ChatGPT, being named is most of the game and there aren't many links to win. On AI Overviews, there are nearly three times as many citation slots, but a mention is scarcer.
Does this match what's already been published?
No, and the disagreement is large enough to be worth stating plainly.
The most specific published figure we found came from Foglift, which measured tools named per answer as: AI Overviews 4.8, Gemini 4.4, Perplexity 2.5, Claude 2.1, ChatGPT 1.0. Our run puts ChatGPT first at 5.28, not last at 1.0, and puts AI Overviews last rather than first. That's not a small gap — it's an inversion of the ranking.
We can't fully explain the difference, and we're not going to pretend otherwise. Two things plausibly contribute: their figure came from 59 answers in a single category against our 491 across ten, and these systems change fast — a measurement taken months apart is measuring a different product. That's a reason to distrust any single snapshot, including ours, rather than a reason to assume we're right.
One claim we deliberately did not test: Wellows reported that two-thirds of AI answers name no brand at all. Our prompts are explicitly commercial — asking for "the best CRM" invites a list of vendors — so we'd expect a far lower no-brand rate by construction, and we measured exactly that. Comparing the two numbers would be dishonest; they have different denominators.
Which engine is most likely to name nobody at all?
Perplexity, by a wide margin. It returned no vendor name at all on 19 of 100 answers, against 3 for Claude, 3 for AI Overviews, 1 for ChatGPT and 0 for Gemini.

That's a specific, actionable asymmetry. On roughly one commercial question in five, Perplexity answers without recommending anyone — it explains the category, links sources and leaves the choice open. No amount of brand optimisation wins a mention that the engine was never going to make. It also fits Perplexity's high citation count: it's the engine most inclined to hand you sources rather than an answer.
What does this change about measuring AI visibility?
It means a single "mentions" number and a single "citations" number describe different halves of the problem, and neither is complete on its own.
If your tracking counts citations, you will systematically under-read ChatGPT, which recommends heavily and links lightly. If it counts mentions, you will under-read AI Overviews, where the citation layer is three times deeper than the naming layer. Either way you get a number that looks like visibility and quietly means something different per engine.
The practical version: track both, per engine, and never average them into one figure. We published the full overlap data showing the same engines share as little as 4.8% of their cited sources — so cross-engine averaging was already suspect, and this adds a second reason. It's also why we're wary of any single cross-engine score, including the one our own product shows.
What do people get wrong about this?
Assuming "getting cited" and "getting recommended" are the same goal. They're measurably not. On ChatGPT you mostly want to be named; on AI Overviews there are far more citation slots than recommendation slots.
Reading one engine's behaviour as the category's. ChatGPT and AI Overviews sit at opposite ends of every measurement here. Whichever you checked first is the one you'll wrongly generalise from.
Treating a published benchmark as settled. We inverted one on our own data. Any figure without a date, a sample size and a method — including this one in six months — should be treated as provisional.
Optimising for engines that weren't going to recommend anyone. One Perplexity commercial answer in five names no vendor at all. That's a ceiling, not a failure of your content.
How did we measure this?
Date: 2 August 2026. Sample: 100 unique buyer-intent prompts across 10 verticals, 10 each. Answers: 491.
Brand counting. We matched answer text against a hand-curated list of 223 real vendors, scoped per vertical, case-insensitive and whole-word. The list is curated rather than derived from the cited domains on purpose: a domain-derived list can only find brands somebody happened to cite, which biases the count toward engines that cite more, and it wrongly counts publishers like TechRadar, PCMag and G2 — and aggregators like Reddit and YouTube — as brands. A fixed list also means a vendor nobody named is still measurable as absent. Names that are ordinary English words (Front, Motion, Close, Monday) were matched only in a disambiguating form, and one — "Ones" — was dropped as unmeasurable.
Engines and surfaces. ChatGPT, Perplexity and Gemini were captured from their consumer interfaces; Google AI Overviews through a SERP data provider; Claude through the Anthropic API's web search tool, because no consumer-interface capture exists for Claude. Claude's numbers therefore describe the API surface. Claude ran on Sonnet 5.
Limitations. 100 prompts is a small sample next to the largest published studies, which run into the hundreds of thousands of answers — treat the overall pattern as far more reliable than any single cell. AI Overviews answered 91 of 100; the nine misses are queries where Google rendered no AI Overview at all, and they're excluded from both numerator and denominator rather than counted as zeros. And this is one day's snapshot of systems that change monthly.
The data. Both halves are published. ai-brands-named-2026-08-02.csv has one row per prompt, engine and brand named — every count in this article recomputes from it. ai-citation-overlap-2026-08-02.csv (JSON) has the citation side. The vendor list itself is in our repository as vendor-vocabulary.json so you can see exactly which 223 names were counted and argue with the choices. We re-run this quarterly.
What should you do with this?
Check your own brand separately on each engine, and check both things: whether you get named, and whether you get cited. If you only ever look at one, you're measuring a different thing on every engine without realising it.
That per-engine split is what GetIntel tracks — see where you stand, or read our honest roundup of AI visibility tools, competitors included, if you're still choosing one.
