
An AI visibility tool that cannot report “I do not know” will report a number instead, and that number will be indistinguishable from a real result.
The four zeros
A visibility score of zero can mean at least four different things.
Genuine absence. The brand is not appearing in answers for those prompts.
A category error. The brand is being scored inside an industry it does not operate in, so the zero is accurate about a market it has never entered.
A retrieval failure. Nothing relevant was retrieved on those runs, which is a property of the run rather than of the brand.
An instrumentation failure. The request errored and was silently scored as a non-citation.
The fourth one is the most dangerous because it is invisible. A practitioner running a fixed question set across five endpoints described exactly this: one endpoint errored on all 32 of its requests, the denominator counted those 32 attempts, and the headline percentage did not move. Half the apparatus was down and the number looked stable.
A number that looks stable while the instrument is broken is worse than a number that drops.
Why “unknown” is a feature and not a defect
The instinct in product design is that unknown looks like weakness. A dashboard full of question marks does not sell as well as a dashboard full of figures.
But consider what the alternative requires. A metric that must always produce a value will produce one for a real result, for a category error, for a retrieval miss and for an outage, and every one of them will render the same way. The confident number is not more informative than the honest gap. It is less informative, because it has destroyed the distinction.
We ship a field that regularly returns not determined. It records whether a cited source is owned by the brand or independent of it, alongside how that call was made: domain match, name inference, or not established. On real scans it lands on inferred or not established more often than we would like.
We show it anyway. Unknown is a real state, and concealing it would be worse than reporting it.
What a report should carry alongside a number
Four things, none of which are expensive to implement.
Attempted and errored counts next to every percentage. This separates small real movement from small silent failure, and small is where most people actually act.
How each determination was made. Whether a classification came from a direct match, an inference, or a default matters as much as the classification.
What is being measured and what is not. A prompt set is a sample. Say what is in it.
An explicit not comparable state. When two engines have executed different tasks on the same prompt, the honest output is that they cannot be compared, not a number that averages them.
The state nobody implements: not comparable
There is one output state that almost no tool offers and it costs nothing to add.
When two engines are compared on the same prompt, the assumption is that they answered the same question. Sometimes they did not. An ambiguous buyer question can be read by one engine as a request to explain a concept and by another as a request to diagnose a situation, and both answers will be perfectly coherent.
Averaging those two outputs into a comparison produces a number. The number means nothing, because it is comparing an explanation with a diagnosis.
The honest output there is not a low score. It is not comparable, with the reason attached. A zero in that slot reads as brand absence and it actually means absence of a shared question, which is a completely different thing to act on.
Checking for it costs one run per engine, before the comparison rather than after.
The harder version of the problem
Some things cannot be determined at all.
How a source came to exist is one of them. Earned, pitched, contributed, placed, sponsored or paid are six meaningfully different states, and most of them cannot be established from outside. For several, nobody except the publication and the client has the information.
An honest field for that would return unknown a great deal of the time. Whether what remains is useful enough to put in front of somebody is a genuinely open question, and we are working through it rather than claiming to have solved it.
The thing we are confident about is the direction. A measurement layer that hides its own uncertainty is asking the market to take its word for it, and that is the position we tell brands not to be in.
How to test a vendor on this in five minutes
Ask three questions.
What does your tool do when a request fails? If failures are scored as non-citations, the trend line is partly measuring their uptime.
Can this metric return an unknown state? If it cannot, ask what it returns when it does not know.
How was this specific classification determined? A vendor that can answer that for one row can answer it for all of them.
FAQs
Why should an AI visibility score be able to return unknown?
Because a metric that can only return a number will return one for a real result, a category error, a retrieval miss and an outage, and all four will look identical in the report.
What is an instrumentation failure in AI visibility tracking?
It occurs when a request to an AI engine errors and the failed attempt is counted as a non-citation rather than excluded, which quietly drags the reported figure down without any change in the brand’s actual position.
How can I tell whether a zero score is real?
Read the competitor list and the attempted and errored counts. A zero accompanied by competitors from an unrelated industry usually indicates a category error rather than absence.
What should a report show alongside a percentage?
Attempted and errored counts, how each determination was made, what the prompt set covers, and an explicit not-comparable state when outputs are not equivalent.
Does showing uncertainty make a tool less useful?
The opposite. Concealing uncertainty removes the distinction between four different situations that require four different responses.
Can a tool determine whether coverage was paid or earned?
Not reliably from outside. Several provenance states are known only to the publication and the client, so an honest provenance field would return unknown frequently.
ABOUT AXIS SUITE
Axis Suite by TrendAxis is the independent intelligence layer that explains what AI believes about your brand, why it believes it, and what decision that belief ultimately drives.
Most tools count citations. Axis Suite tells you which of those citations are actually evidence about you, how many separate parties they represent, and what would remain if your largest source disappeared.
Run a scan at axissuite.ai or read the methodology in the Proof Center.