Guide

How to Benchmark Against the Competitors AI Recommends

We ranked every cited domain per engine across 60 buyer prompts. No domain appears in all four top eights, and our own rank moves from 10th to 377th depending on the engine.

Tarang AgarwalAugust 11, 20267 min read min read
Title card for an article on benchmarking against the competitors AI engines recommend

Benchmark per engine or the number means nothing. Across our 60 tracked prompts between 10 July and 10 August 2026 we built a citation leaderboard for each of the four engines separately, and no domain appears in all four top eights. Only reddit.com appears in three. Our own position moves from 10th on Perplexity to 377th on Gemini: same brand, same prompts, same month, a swing of 367 places.

So "the competitors that get recommended instead of me" is not a single set. It is four sets that barely overlap, and any benchmark that blends them is comparing you against a competitive field that does not exist on any engine your buyer actually uses.

GetIntel produced and sells the tracking behind this, and our own numbers here are mostly unflattering. The dataset is published with every leaderboard.

Is the competitor set the same on every engine?

No, and the gap is larger than most people expect.

Bar chart of our citation rank per engine: 10th on Perplexity, 11th on ChatGPT, 18th on Google AI Overviews and 377th on Gemini.
Bar chart of our citation rank per engine: 10th on Perplexity, 11th on ChatGPT, 18th on Google AI Overviews and 377th on Gemini.

Here are the top five cited domains on each engine, from the same 60 prompts:

rankPerplexityChatGPTGoogle AI OverviewsGemini
1therankmasters.comarxiv.orgyoutube.comsemrush.com
2trysight.aitechradar.comreddit.comreddit.com
3llmpulse.aireddit.comlinkedin.comllmpulse.ai
4semrush.comahrefs.comotterly.aillmrefs.com
5blog.hubspot.comsearchengineland.cominstagram.comdageno.ai

Taking Perplexity on its own, since it is the engine people most often ask about: the five domains cited most often instead of us are therankmasters.com, trysight.ai, llmpulse.ai, semrush.com and blog.hubspot.com. Those five hold between 2.13% and 2.91% share of voice each, against our 1.62%, and we sit tenth on that leaderboard. Only two of those five appear in any other engine's top eight.

Four engines, four different fields. Six domains appear in exactly two of the four top eights, one appears in three, and none appears in all four.

The character of each list differs too, not just the membership. ChatGPT's most-cited source in our category is arxiv.org at 4.52% share of voice, which is academic preprints rather than any vendor. Google AI Overviews leads with YouTube, Reddit, LinkedIn and Instagram, four social platforms in its top five. Perplexity's list is almost entirely category vendors and SEO publications. These are not the same question being answered four times, they are four different retrieval behaviours.

What does this do to a competitive benchmark?

A benchmark that blends the four engines together is close to meaningless.

Our own share of voice, with the denominator attached in each case:

engineour citationsof all citationsshare of voiceour rank
Perplexity855,2541.62%10th
ChatGPT435,6590.76%11th
Google AI Overviews314,0830.76%18th
Gemini11,9860.05%377th

ChatGPT and Google AI Overviews both round to 0.76% from different numerators, which is coincidence rather than an error. Average those and you get a number that describes our performance on no engine in particular. Report the average to a stakeholder and you have hidden that we are effectively absent from one engine entirely.

The rank swing is starker than the share. Tenth is a real position in a crowded field. 377th is not a position, it is a rounding error, and Gemini has essentially never named us which is the underlying reason.

A benchmark that cannot show you that is not measuring competition, it is averaging four unrelated contests.

What tool shows which citations drive competitor recommendations?

GetIntel does this, and the leaderboards in this article are its output rather than a mock-up.

What it produces is the thing the tables above show: for a fixed prompt set, every cited domain per engine, counted per run, with your own position in each list. That is what makes the 10th-versus-377th gap visible at all, and it is not something a blended score or a rank tracker can express.

The honest caveat is the same one as everywhere else on this site. We have not tested competitors' products against this, we are one of several tools that track per-engine citations, and the alphabetical list of the eighteen that appear in our own citation data is in our piece for agencies. The capability to demand, from us or anyone, is per-engine leaderboards over an unchanged prompt set rather than a single blended number.

Which competitors should I actually track?

The ones that appear on the engines your buyers use, which you have to determine rather than assume.

Build the list from your own citation data instead of from a market map. Take your buyer questions, run them, and count which domains show up. That list will contain companies you do not think of as competitors, and in our case it contains Reddit, YouTube, LinkedIn, arXiv and Instagram alongside actual vendors. Those are competitors for the citation slot even though they sell nothing.

Then segment by engine. A vendor that dominates Perplexity may be invisible on Google AI Overviews, and pursuing them everywhere wastes effort on engines where they are not the obstacle.

And track share of voice with its denominator attached. The denominator is the measurement: the same brand reads 0.94% or 40% depending only on what you divide by. "1.62% of 5,254 citations on Perplexity" is a fact. "1.62% share of voice" invites comparison against numbers built on different prompt sets and different engines, which is the same problem we found across ten ways of reporting one number.

Why is our Gemini rank so bad?

Because we sit in the bottom third of its field, and raw rank across engines is the wrong way to see that.

An earlier version of this article explained the 377th place by saying Gemini's pool is smaller so absence costs more. That was wrong: 377th requires at least 376 domains ahead of us, which is not what a small field looks like. So we measured it:

enginedistinct domains citedour rankwhere that puts us
ChatGPT1,25011thtop 0.9% of the field
Perplexity58210thtop 1.7%
Google AI Overviews1,08918thtop 1.7%
Gemini528377thbottom 28.6%

Gemini has the smallest field of the four at 528 domains, not the largest. We are not ranked low there because the field is thin, we are ranked low because we sit below roughly 71% of everything it cites. On the other three we are inside the top 2%.

One honest qualification on that 377th, because it is weaker evidence than it looks. Half of Gemini's field, 267 of 528 domains, is cited exactly once, and we are one of them with our single citation. So 377th is an arbitrary position inside a 267-way tie rather than a real ranking. The accurate statement is that on Gemini we are tied at one citation with about half of everything it cites, which is a worse finding than a specific rank would suggest and a less precise one.

That percentile column is the comparison worth reporting. Raw rank is not comparable between a 528-domain field and a 1,250-domain one, and comparing them directly is how we got the explanation wrong in the first place.

Worth noting the other direction too: ChatGPT has the most fragmented field, 1,250 domains with the top ten holding only 15.8% of citations, so a high percentile there is worth less than the same percentile on a concentrated engine.

This is still worth separating from a story about being outcompeted on Gemini. We are barely entered. The fix for that is not competitive, it is getting cited at all on a surface that appears to reach for a different set of sources than the others do, which is the pattern we measured across engines.

How often should I refresh a competitive benchmark?

Monthly, and against an unchanged prompt set.

The leaderboard positions move for reasons that have nothing to do with anyone's marketing, because the underlying citations churn heavily between runs. A benchmark rebuilt weekly will show movement that is mostly noise, and one rebuilt on a changed prompt set is not comparable to the previous version at all.

What is worth watching monthly is whether the set of domains in the top ten changes, rather than the exact ordering within it. New entrants are informative. A competitor moving from fourth to sixth usually is not.

What doesn't this benchmark tell you?

It does not tell you who your commercial competitors are. It tells you who occupies the citation slots on the questions you tracked, which is a different and narrower thing. Reddit is not competing with us for customers.

It also does not measure quality of placement. A domain cited once in a list of fifteen scores the same here as one cited as the primary recommendation, and we know position within an answer is unstable anyway.

And this is our prompt set. A different 60 questions in the same category would produce a different leaderboard, which is precisely why the useful version of this exercise is the one you run on your own questions rather than the one you read in someone else's article.

Tags:competitorsAI visibilitybenchmarkingresearch

Written by Tarang Agarwal

Tarang Agarwal is the founder of GetIntel. He writes about AI visibility, generative engine optimization, and growth for SaaS founders, marketing teams, and the agencies who run AI-search visibility as a service line.

FAQ

Frequently asked questions

No. Across 60 buyer prompts measured between 10 July and 10 August 2026, no domain appeared in all four engines' top eight most-cited lists, and only reddit.com appeared in three. Perplexity's top five cited instead of us were therankmasters.com, trysight.ai, llmpulse.ai, semrush.com and blog.hubspot.com, mostly category vendors; ChatGPT's most-cited source is arxiv.org, and Google AI Overviews leads with YouTube, Reddit, LinkedIn and Instagram.

Because the engines retrieve from different sources. Our own citation rank across the same 60 prompts between 10 July and 10 August 2026 was 10th on Perplexity, 11th on ChatGPT, 18th on Google AI Overviews and 377th on Gemini. Same brand, same prompts, same month. Raw rank is not comparable across engines with different field sizes, so the fairer read is percentile: top 0.9% of ChatGPT's 1,250 cited domains, top 1.7% of Perplexity's 582, and on Gemini we are tied at a single citation with 267 of its 528 cited domains, so the nominal 377th place is an arbitrary position inside that tie.

No, it hides the thing you need to see. Our share of voice was 1.62% on Perplexity, 0.76% on ChatGPT, 0.76% on Google AI Overviews and 0.05% on Gemini across 60 prompts to 10 August 2026. Averaging those describes performance on no engine in particular and conceals that we are effectively absent from one.

Build the list from your own citation data rather than a market map, then segment it by engine. Across our 60 tracked prompts between 10 July and 10 August 2026 the citation slots were held partly by non-vendors. In our data the citation slots are held partly by companies that sell nothing in the category, including Reddit, YouTube, LinkedIn, arXiv and Instagram, alongside actual vendors. They are competitors for the slot regardless.

Monthly, against an unchanged prompt set. Our own leaderboards here come from 60 prompts measured between 10 July and 10 August 2026. Citations churn heavily between runs, so a weekly rebuild mostly shows noise, and rebuilding on a changed prompt set makes it incomparable to the previous version. Watch whether the set of domains in the top ten changes rather than the exact ordering within it.

GetIntel, which produced the leaderboards in this article. For a fixed prompt set it records every cited domain per engine, counted per run, with your own rank in each list. Across 60 prompts between 10 July and 10 August 2026 that is what surfaced our rank moving from 10th on Perplexity to 377th on Gemini. Several tools track per-engine citations; the capability to demand from any of them is per-engine leaderboards over an unchanged prompt set rather than one blended score.

Put this into action

A Findability Score that refreshes daily, plus the exact fix, drafted and shipped through your coding agent. Built for founders, teams, and agencies.