Guide

How Many Reruns Before an AI Visibility Number Means Anything?

On prompts whose answer actually varies, a single check is off by 33 percentage points on average. Here is what we measured across 320 prompt-and-engine series.

Tarang AgarwalAugust 11, 20266 min read min read
Title card for an article measuring how many reruns an AI visibility number needs before it is reliable

It depends entirely on whether the prompt is one of the stable ones, and the uncomfortable part is that you cannot tell in advance. Across 320 prompt-and-engine series with at least ten runs each, 275 never changed their answer at all. On those, one check is the whole truth. On the 45 that did vary, a single check was off by 32.7 percentage points on average and landed within 10 points of the real rate only 31.1% of the time.

That is the honest shape of the answer. Most of your prompts need one run. A minority need closer to ten. Nothing tells you which is which except having run them.

Everything below comes from our own tracking between 10 July and 10 August 2026, four engines, 11 to 12 runs per series. GetIntel sells a tracker that automates this, which is a reason to check the numbers rather than take them: the dataset is published.

Why can one run be so wrong?

Because a single run can only return 0% or 100%, and the truth is usually neither.

If a prompt names you in roughly half its answers, a single check reports either 0 or 100 and both are wrong by about 50 points. The error is not noise around a good estimate. It is a category error: you are reading a coin flip as a measurement.

That is why the average miss on varying prompts is 32.7 points rather than something small. It is not that the engine is unstable in some exotic way. It is that one observation of a probabilistic process is not an estimate of its rate.

How much does each extra run buy you?

Bar chart. Mean error of the cited-rate estimate falls from 32.7 percentage points at one run to 24.0 at three, 18.5 at five and 12.3 at seven runs.
Bar chart. Mean error of the cited-rate estimate falls from 32.7 percentage points at one run to 24.0 at three, 18.5 at five and 12.3 at seven runs.

Measured on the 45 varying series, error against each series' own final rate:

runsmean errorwithin 10 pointswithin 20 points
132.7 pts31.1%44.4%
324.0 pts31.1%48.9%
518.5 pts40.0%62.2%
712.3 pts55.6%86.7%

The practical read is that the first few runs buy you very little and the gains arrive later than intuition suggests. Three runs still misses by 24 points on average, which is worse than most people assume when they check something "a few times to be sure". Seven runs brings the error to 12.3 points, a little over a third of where it started.

So is the answer ten runs?

Our series only run 11 to 12 deep, and that ceiling contaminates the tail of the table.

Each estimate is scored against that series' own final rate, which is itself computed from 11 or 12 runs, so the two share most of their data and of course they agree. The bound is exact: for a series of length N, an estimate after k runs cannot differ from the final rate by more than (N-k)/N. At N=11 and k=10 that ceiling is 9.1 points, so "100% within 10 points" at ten runs is guaranteed by arithmetic before any engine is queried. It says nothing about ChatGPT and everything about our series length.

The same bound bites earlier than we first thought. At eight runs on an 11-run series the ceiling is 27 points, which is loose enough to matter, so we have cut the table at seven runs. The 8, 9 and 10-run rows remain in the published dataset, correctly labelled, rather than in the argument.

So read the trend from one to seven runs, which is not contaminated that way, and treat anything past it as an artefact of our ceiling. We could tell you "ten runs gets you within 10 points 100% of the time" and it would be a true sentence about our data and a misleading sentence about yours.

The defensible claim is narrower: seven runs cuts the error from 32.7 points to 12.3, a reduction of about 62%, and we cannot say from this data where it flattens.

Do most prompts actually change their answer?

No. Of 320 series, 275 never varied at all, which is 85.9% and the part that changes what you should actually do.

For those prompts, the tenth run tells you exactly what the first one did. All the rerunning in the world adds nothing, and a vendor charging you per run on those is selling you repetition rather than information.

Which means the real value of a schedule is not that it averages away noise on every prompt. It is that it identifies the minority where noise exists at all. You are not buying precision across the board. You are buying the ability to find your 14%.

This also explains a common and reasonable complaint: people run daily tracking, watch a number sit perfectly still for weeks, and conclude the tool is broken. It usually is not. Most prompts genuinely do not move, and the source lists underneath them churn almost completely while the brand answer stays put.

What should I actually do?

Four things, roughly in order of value.

  • Never act on a single check of a prompt you have not run before. If it comes back without your name, that is one draw. On a varying prompt a single check missed the real rate by more than 10 points 68.9% of the time.
  • Run a burst before you run a schedule. Five to seven runs of a new prompt set over a few days tells you which prompts vary. That is a much better use of budget than a daily cadence applied uniformly forever.
  • Then reduce cadence on the stable ones. If a prompt has returned the same answer ten times, checking it daily is spending money to be told the same thing.
  • Report a rate with its run count attached. "Cited on 3 of 11 runs" is a fact. "27% visibility" is that same fact with the evidence removed, and it invites comparison against numbers built on one run, or against numbers built on a different retrieval path entirely.

How do I tell if a drop is real or just noise?

Compare it against how much the number moves when nothing has happened.

We ran that test on ourselves. Splitting each series into its first and second half, with no publishing or campaign in between, 13.1% of all 320 series showed a swing of 10 points or more purely from noise. On the 45 series that vary at all, 93.3% did.

Bar chart. Splitting series into halves with nothing happening in between, 13.1% of all 320 series swung 10 points or more and 8.1% swung 20 or more; among the 45 varying series those figures are 93.3% and 57.8%.
Bar chart. Splitting series into halves with nothing happening in between, 13.1% of all 320 series swung 10 points or more and 8.1% swung 20 or more; among the 45 varying series those figures are 93.3% and 57.8%.

So on a prompt you already know to be unstable, a 10-point drop between two periods is close to meaningless by itself. On a prompt that has never varied, the same 10-point drop is worth investigating immediately, because that prompt has no history of moving.

One honest wrinkle: the moves were not symmetric. At the 10-point threshold, 36 series drifted up and only 6 drifted down. Some of that upward drift is real rather than noise, because our citation count genuinely grew over the window, which means this test overstates noise upward and the downward figures are the cleaner read of it.

What does this not establish?

One brand, one category, four engines, 45 varying series. That is a small group, and 45 series is enough to see a clear direction but not to pin a threshold.

We also measured a binary, whether the brand was cited at all. A rank or share-of-voice metric has more states and will almost certainly need more runs than this to stabilise, not fewer.

And the split between stable and varying prompts is a property of our prompt set and our category, not a law. A more contested category would likely have more varying prompts, which would raise the value of reruns and change the arithmetic above.

If you want the version that applies to you rather than to us, run your own prompts five to seven times over a week and count how many gave a different answer. That number is the only one that decides your cadence, and it takes a few days to get.

Tags:AI visibilitymethodologymeasurementresearch

Written by Tarang Agarwal

Tarang Agarwal is the founder of GetIntel. He writes about AI visibility, generative engine optimization, and growth for SaaS founders, marketing teams, and the agencies who run AI-search visibility as a service line.

FAQ

Frequently asked questions

Run a burst of five to seven on any prompt you have not run before, then drop to a single check on the roughly 86% that never vary. Across 320 prompt-and-engine series measured 10 July to 10 August 2026, 275 never changed state at all. On the 45 that did, seven runs cut the mean error from 32.7 percentage points to 12.3.

On a prompt that actually varies, a single check missed the true rate by 32.7 percentage points on average and landed within 10 points only 31.1% of the time, across 45 varying series to 10 August 2026. A single run can only return 0% or 100%, so on any prompt with a middling rate it is wrong by construction.

No. Across 320 prompt-and-engine series with at least ten runs each, measured 10 July to 10 August 2026, 275 never changed state at all. That is 85.9%. The value of repeated runs is not averaging noise everywhere, it is identifying the roughly 14% of prompts where noise exists.

We will not claim that. Across 320 series measured 10 July to 10 August 2026, our series run only 11 to 12 deep, so at nine or ten runs the estimate shares most of its data with the final rate it is scored against, making the convergence there partly definitional. The uncontaminated part of our data is one through seven runs; for a series of length N an estimate after k runs cannot differ from the final rate by more than (N-k)/N, which at N=11 and k=10 is only 9.1 points, and it shows error falling to from 32.7 points to 12.3 by seven, a reduction of about 62%.

With its run count attached, as in our own 320-series measurement of 10 July to 10 August 2026. "Cited on 3 of 11 runs" carries its own evidence; "27% visibility" is the same fact with the sample size removed, which makes it look comparable to a number built from a single check when it is not.

Compare it to how much the number moves when nothing has changed. Splitting 320 prompt-and-engine series into halves with no activity in between, 13.1% swung by 10 points or more on noise alone, and among the 45 series that vary at all, 93.3% did. A 10-point drop on a historically unstable prompt is close to meaningless; the same drop on a prompt that has never varied is worth investigating. Measured 10 July to 10 August 2026.

Put this into action

A Findability Score that refreshes daily, plus the exact fix, drafted and shipped through your coding agent. Built for founders, teams, and agencies.