
One run of a prompt tells you almost nothing, and the number of runs it takes to tell you something depends entirely on which question you are asking.
The IAB put this plainly in its August guidance: single-response measurement is not measurement.
Why one run is an anecdote
Run the identical buying question against ChatGPT several times in one sitting. Not variations of the question. The identical string.
You will not get the same answer. You will get a distribution.
A single value implies a stability that does not exist. A screenshot of Perplexity recommending your brand is real, it happened, and it is the least demanding condition your brand will ever be measured under.
Two brands, one average, opposite situations
Consider a made-up example to make the arithmetic visible.
Brand A appears in nine runs out of ten, usually low in the response. Brand B appears in five out of ten, near the top every time it appears. Depending on how a tool weights position against appearance rate, those two can land on a similar composite score.
Which means the composite is trading appearance rate against prominence at roughly two to one, and nobody agreed to that exchange rate. It is buried in the scoring.
The two situations are not remotely alike. Brand A has support the engine reaches for almost every time. Brand B sits on a threshold, present when retrieval happens to surface a supporting source and absent when it does not. One new competitor comparison article and Brand B disappears from that query while Brand A barely moves.
Brand B is not less visible on average. It is less settled.
What a given number of runs can actually support
This is where most testing advice stops, and it is the part that decides whether your finding survives scrutiny.
Using exact binomial confidence intervals, ten runs is close to useless for establishing stability. A brand appearing in eight of ten runs has a true appearance rate somewhere between roughly 44 and 97 percent. That range is too wide to act on.
Twenty runs is better and still asymmetric:
Appearing in twenty of twenty runs supports a true rate of roughly 83 to 100 percent, which is a genuine stability finding.
Appearing in sixteen of twenty gives roughly 56 to 94 percent, which cannot support a claim of eighty percent stability.
Appearing in eleven of twenty gives roughly 32 to 77 percent, which is genuinely unstable and is a real finding.
The practical consequence: twenty runs can demonstrate instability and can barely demonstrate stability. Expect most cells to come back inconclusive, and treat that as the honest reading rather than a failed test.
Thirty runs tightens things meaningfully. Twenty-four of thirty gives roughly 61 to 92 percent. If one query comes back interesting and inconclusive, take that query to thirty rather than taking everything to thirty.
The setup that decides whether any of it counts
Turn memory off. A new ChatGPT chat does not clear memory. Use a temporary chat with memory disabled, or run logged out. This is the single most common way a test like this gets quietly contaminated, and it invalidates everything downstream.
Keep the string identical. Not even punctuation varies.
Run one prompt and one engine in a single sitting. An engine update between runs is a confound you cannot see.
Hold browser, location and account state constant. Otherwise personalization variance is mixed into your result.
Test the consumer product, not the API. The API is a different system with different retrieval behavior. If the question is about what buyers see, the consumer surface is the right instrument.
Repeat on a second engine. Claude, Gemini and Perplexity behave differently from ChatGPT, and differences between them are only interpretable once you know how much a single engine varies with itself.
The question nobody has answered
Every cross-engine comparison currently being designed rests on an assumption: that when you ask an engine the identical question repeatedly, it performs the same kind of task each time.
If an engine explains the category on some runs and produces a shortlist on others, then the task itself is a distribution, and comparing two engines is comparing two draws rather than two engines.
Nobody has published that number. We are measuring it: three prompts, two engines, twenty runs each, with the method written down before the first run and the responses labeled blind afterward.
What to record on every run
The fields matter as much as the count, because a run you cannot re-examine later is a run you have to take on trust.
Record the run number, the engine, the exact prompt, the timestamp, whether the brand appeared, roughly where in the response it appeared, and which sources were cited.
Then record the full raw response text. This is the one people skip and the one that matters most. Any set of parsed fields you design today will turn out to be missing something you need in three months, and the raw response is the only artifact that lets you go back and ask a question you had not thought of yet.
Deciding in advance what you will conclude, including what a boring result means, is the discipline that stops a null becoming “the measurement was flawed” three months later.
FAQ
Is ten runs of a prompt enough?
Not for a stability claim. Appearing in eight of ten runs corresponds to a true rate anywhere between roughly 44 and 97 percent, which is too wide a range to act on. Ten runs can tell you whether you are looking at a coin flip, which is still more than one run tells you.
Why does the same prompt give different answers?
AI responses are probabilistic and retrieval varies between runs. What gets pulled into the context window differs, which changes which brands are available to name in the answer.
Does a new chat clear ChatGPT’s memory?
No. Memory persists across chats unless it is explicitly disabled. Use a temporary chat or run logged out, otherwise your results reflect your own history rather than a shared baseline.
Should I test ChatGPT, Claude, Gemini and Perplexity the same number of times?
Yes, if you intend to compare them. Unequal run counts make differences between engines uninterpretable, because you cannot tell an engine difference from a sampling difference.
What is decision-grade measurement?
It is the IAB’s term for measurement robust enough to move budget against, requiring large and diverse query sets, frequent testing, reproducibility, disclosed variation ranges and stated confidence levels. Lighter measurement is described as directional and is suitable for trend spotting rather than decisions.
How often should this testing be repeated?
The IAB guidance suggests monthly or quarterly for directional purposes and weekly or more frequent for decision-grade. What matters more than the interval is that it stays constant, since a changing cadence makes trends uninterpretable.
ABOUT AXIS SUITE
Axis Suite by TrendAxis is the independent intelligence layer that explains what AI believes about your brand, why it believes it, and what decision that belief ultimately drives.
Most tools count citations. Axis Suite tells you which of those citations are actually evidence about you, how many separate parties they represent, and what would remain if your largest source disappeared.
Run a scan at axissuite.ai or read the methodology in the Proof Center.