Judge tools on whether they let you defend a number, not on how many engines they list. We took one brand's data for one month and computed the headline visibility percentage ten defensible ways, changing nothing but the definition, the engine selection and the window. It came out anywhere from 0% to 40%. Every one of those numbers is honest, and a client shown any single one of them has no way to know which choices produced it.
That is the central problem of agency reporting in this category, and it is not solved by picking a better dashboard. It is solved by fixing your definitions before you have a result you like.
The figures below are our own brand across 60 prompts, four engines, 10 July to 10 August 2026. GetIntel is one of the tools you would be choosing between, so weigh the recommendation accordingly. The dataset is published.
How much does the reporting choice actually move the number?
By more than most client work ever will. All ten figures below come from the same 60 prompts across four engines, 10 July to 10 August 2026.

| how it is reported | result |
|---|---|
| Any run names us, all four engines | 40.0% |
| Any run, last 7 days only | 36.7% |
| Any run, ChatGPT only | 23.3% |
| Any run, Google AI Overviews only | 21.7% |
| Any run, Perplexity only | 18.3% |
| Share of all runs naming us | 7.0% |
| Any run, Gemini only | 1.7% |
| Named in a majority of its runs | 0.0% |
Same brand, same prompts, same runs, same month. The spread is 40 points. The table shows eight rows against ten definitions in the dataset: "named in every run" also returns 0.0% and "excluding Gemini" also returns 40.0%, so both are folded into the rows they duplicate.
The biggest single lever is the definition of visible. Counting a prompt as covered when any run names you gives 40%. Requiring a majority of runs gives zero, because no prompt in our set names us in more than half its runs. Our citations are real and they are sporadic, and those two facts produce wildly different headlines depending on which one your tool encodes.
Which number should go in the client report?
Whichever one you can still defend in month six, chosen before you see it.
The practical answer is to write your definition down at the start of an engagement and keep it fixed: which engines count, how many runs, what threshold makes a prompt "covered", what window. Then report the same way every month even when a different choice would look better. An agency that switches from all-engine to ChatGPT-only reporting between months has shown a client a 16.7-point move that no work produced.
Our own preference, for what it is worth, is to report a count rather than a percentage. "Named on 24 of 60 tracked prompts" carries its own denominator and cannot be quietly restated. A percentage strips the sample size out and invites comparison against a competitor's number built on different rules.
And exclude no engine silently. Gemini reads 1.7% for us against ChatGPT's 23.3%, so dropping it lifts the headline from 36.7% to 40.0% without anything changing in the world. If you exclude an engine, say so and say why in the same sentence as the number.
What should I demand from the tool itself?
Five things, and only the first two are about features.
- Configurable and visible definitions. If you cannot see how the tool defines a covered prompt, you cannot defend the number to a client who asks. This is the single most important thing on the list and almost nobody evaluates it.
- A data model built for answers, not for rankings. Position is not stable in an AI answer: a surviving source keeps its place only 39.5% of the time. A tool storing a score and a position has fitted the new problem into the old schema.
- Per-engine breakouts, not just a blended score. A blended number hides a 21.6-point spread between engines in our data, from 23.3% on ChatGPT to 1.7% on Gemini. Clients in different categories will have different weak engines, and you cannot advise on that from an average.
- Raw answer text per run. When a client asks why a competitor is named, the answer lives in the response, not in the score. Without it you are guessing.
- Run counts on every figure. A single check on a prompt that varies is off by 33 percentage points on average, so a number without a run count behind it is not comparable to one with.
- Per-client isolation you can actually verify. Multi-brand tools differ enormously here, and the failure mode is quiet: one client's data appearing in another's report is worse than no report. In fairness, this is the one item on the list we are recommending from reasoning rather than measurement. We have not tested any tool's isolation, including our own, in a way we could publish.
Which tools should an agency actually evaluate?
The honest shortlist is the one the engines themselves surface. Rather than give you our opinion, here are the 18 tools in this category that appear among the 50 most-cited domains in our own probe corpus, with how often each was cited between 10 July and 10 August 2026. We are on the list, sixteenth of eighteen by citation count, which is the main reason to trust the list at all.
Listed alphabetically, deliberately, so nothing here reads as a ranking. Ahrefs (442), AI Clicks (381), dageno.ai (695), Foglift (203), Frase (402), GetIntel (276), Kime (507), LLM Pulse (580), LLMrefs (360), Otterly (521), Profound (499), Rankability (203), SE Ranking (393), Semrush (659), Siftly (286), The Rank Masters (526), Trysight (543), Useomnia (355).
Read that as a candidate list, not a ranking. Citation count measures how often a domain is cited on AI-visibility questions, which reflects how much it publishes about the category and how well that content is retrieved. It does not measure product quality, and Semrush and Ahrefs are general SEO suites that also ship AI-visibility features rather than dedicated tools. Our own 276 puts us sixteenth of the eighteen.
We have not tested any of these against the five criteria above, including our own, and we are not going to publish claims about competitors' features that we have not verified. Take the five questions to each vendor directly. Any of them that cannot show you a configurable definition and raw answer text should be easy to eliminate in one call.
Is there a tool that does per-client reporting for multiple brands?
Yes. Several tools in the list above are built for multi-brand use, GetIntel among them, and per-client reporting is a normal feature rather than an exotic one. The part worth checking is not whether a tool offers per-client reports but whether the definitions behind those reports are visible and fixed, because that is what determines whether the number you hand a client in month one is the same kind of number you hand them in month six. The governance side of that, keeping prompts and competitor sets comparable across sub-brands, is covered separately in AI visibility monitoring for multi-brand portfolios.
Is my client's number normal?
Probably not, because there is barely a normal to be near. Across the 63 brands on our platform with at least 20 probe runs, 58,016 runs in total to 11 August 2026, the distribution of how many of a brand's own prompts name it is strongly bimodal.

| where a brand sits | brands |
|---|---|
| Named on zero of its own prompts | 14 |
| Under 10% | 21 |
| Between 10% and 50% | 16 |
| At or above 50% | 26 |
Before reading anything into those buckets: this is one platform's customer base, which self-selects for companies that already suspect they have a visibility problem, so it is not a random sample of businesses and not a benchmark in the strict sense.
The median is 28.8% and the mean is 41.4%, but neither describes a typical brand well. Only 16 of 63 sit in the 10 to 50% band, which is a quarter of them rather than a vanishing minority, but it is the smallest of the three groups. Most are either largely absent or largely present, and 14 are named on not a single one of their own tracked prompts.
We checked the two obvious explanations and neither accounts for the shape. Prompt-set size does not create it: split at 50 prompts, both groups are bimodal with medians of 28.8% and 30.0%. The two are not identical though, and the difference cuts against us rather than for us - the smaller-prompt-set group is the more polarised of the two, with 30.8% of its brands at zero against 16.2%. Size sharpens the effect without causing it. It is not recency either: brands named on zero prompts have a median of 300 probe runs, exactly the same as every other brand. The invisible ones have been measured just as thoroughly.
For an agency this changes the pitch. A client at 4% is not slightly behind, they are in the larger of two clusters, and moving them is a step change rather than an optimisation. A client at 60% has a retention problem rather than an acquisition one. The middle, where incremental gains are the natural story, holds the fewest brands of the three groups.
Two honest caveats. This is one platform's customer base, which self-selects for companies that already suspect they have a problem, so it is not a random sample of businesses. And prompt sets are chosen per brand, so these are not comparable instruments in the way a standard benchmark would be. Read the shape, not the percentile. The aggregate data is published, with no brand names, industries or per-brand rows in it.
How do I show a client before-and-after data?
Carefully, because the noise is larger than most of the movements you will want to claim. The figures here come from 320 prompt-and-engine series measured 10 July to 10 August 2026.
We measured this by splitting our own tracking into halves with nothing happening in between. On prompts that vary at all, 93.3% swung by 10 points or more on noise alone. So a 10-point improvement in a client report, presented without context, is more likely to be the measurement moving than the market.
What survives scrutiny is a change in the count of prompts where the client is named at all, measured on a fixed prompt set over at least a month, with the run count attached. What does not survive is a percentage that moved between two single checks.
The honest framing to a client is that most prompts do not move. In our own data 275 of 320 prompt-and-engine series never changed state at all. Winning one new prompt is a real result and it is worth reporting as one prompt, not as a percentage that makes it sound like a trend.
Why do my clients keep seeing competitors instead of them?
Often there is no consistent competitor to point at, which is worth knowing before you build a strategy around one.
We checked whether any domain reliably owns a buyer question in our category. Only 27 of 60 prompts had one, spread across 20 different domains, and the biggest single holder was Reddit. The full working is in our piece on diagnosing this.
For a client, that usually reframes the conversation from "beat this rival" to "be present on this question at all", which is a different and more achievable brief.
Does this generalise beyond one brand?
Partly. One brand, one category, 60 prompts, four engines, a single month. The size of the spread will be different for your clients. That a spread exists will not be.
And the ten definitions above are not exhaustive. They are ten reasonable ones. A determined person could construct a more flattering number than 40% or a bleaker one than 0%, which is rather the point.
