Why Your AI Visibility Tools Disagree About Your Brand

Buried in the IAB’s August 2026 measurement guidance is a sentence that should unsettle anyone who has bought an AI visibility tool.

More than 20 companies now sell AI visibility measurement. They use different methodologies. They produce materially different results for the same brand.

Which means you could buy two platforms, scan the same company in the same week, and walk into a board meeting holding two versions of reality.

The reflex response is to wait. Let the standards settle, let the vendors converge, buy something once the numbers agree.

That reflex is understandable and it is wrong, because two completely different kinds of disagreement are happening here. One of them is a problem to be fixed. The other is the most useful signal in your report.

The disagreement that is noise

Most vendor disagreement is methodology.

One tool runs 40 queries a month. Another runs 400 a week. One counts any appearance of your brand name as a mention. Another counts it only when the model attaches a source link. One phrases the query as “best CRM software” and another as “what CRM should I use for a mid-size sales team,” which are different buying conversations that will produce different brands. One samples ChatGPT only. Another blends ChatGPT, Claude, Gemini, and Perplexity into a single score without disclosing the weights.

None of those tools is lying. They are answering different questions and printing the answers in the same typeface.

This is precisely what the IAB guidance is built to address, and its remedy is the right one. Not a mandated methodology, which would freeze the field prematurely, but disclosure. How many queries. Across how many intent types. How often. With what variation range. Reported as a range rather than a single value, because a single value implies a stability that probabilistic systems do not have.

If two tools disagree and neither will explain how it sampled, you do not have a measurement disagreement. You have two unlabeled numbers.

Ask for the labels. That part is solvable this quarter.

The disagreement that is information

Then there is the second kind, and it does not go away when the disclosure standards land.

Same query. Same week. Careful sampling on both sides. ChatGPT recommends you confidently. Perplexity does not mention you at all. Gemini names you but places you in the wrong category. Claude mentions you only when the question includes a specific constraint.

That is not measurement error. That is four systems genuinely holding different beliefs about your brand.

And that gap points somewhere specific. These systems retrieve from different source sets, weight recency differently, and require different amounts of corroboration before moving from mentioning a brand to recommending it. When one engine is confident about you and another has never surfaced you, the difference usually sits in the source layer rather than in your content.

The engine that recommends you is likely reaching a source type the other does not index or does not weight heavily. The engine that ignores you is telling you which part of your evidence footprint is thin.

A brand that shows up strongly in one place and nowhere else does not have a content problem. It has a source concentration problem, and no amount of publishing on its own domain will fix it.

The disagreement is the diagnosis.

A cheap test to tell them apart

You can run this yourself in about twenty minutes without buying anything.

Take one buying question that genuinely matters to your business. Run it by hand across ChatGPT, Claude, Gemini, and Perplexity in a single sitting. Write down what each says about you, which category it places you in, and what language it uses. Then run the same set again the next day.

If your manual runs broadly agree with each other but your two vendors disagree, the gap is methodology. Go ask both vendors for their sampling disclosure and reconcile from there.

If your manual runs disagree in roughly the same pattern your vendors do, the disagreement is real. That is not something to average away. That is where the work is, and it is now pointing you at a specific engine and a specific gap rather than at a vague instruction to publish more.

Three things not to do

Do not pick the tool with the friendliest number. Everyone knows this. A surprising number of teams do it anyway, usually without deciding to.

Do not average four engines into one score and present it without the spread. A brand scoring 60 across all four engines and a brand averaging 60 from scores of 95, 80, 45, and 20 are in completely different strategic positions. One has broad consistent recognition. The other has one strong pocket and three gaps. A single number erases that distinction entirely, and the distinction is the actionable part.

Do not wait for the industry to converge before you start measuring. The standard published this month is a floor, not a finish line. Teams that begin now will have a year of trend data by the time everyone agrees on the vocabulary, and trend data is the only thing that turns a snapshot into a position.

The part worth watching

The industry is about to spend the next year working hard to make these tools agree with each other.

Much of that work is necessary and overdue. Disclosure standards, sampling minimums, reported ranges instead of false precision, clear definitions separating a mention from a citation from a recommendation. All of it improves the field.

But there is a version of that project that goes too far. If the goal becomes a single reconciled number that every vendor reproduces identically, the field will have smoothed away the variance that was telling brands where their evidence gaps are.

Convergence between measurement tools is progress.

Convergence that erases the disagreement between the models themselves is not measurement getting better. It is measurement getting quieter.

When your tools disagree about your brand, the first question is not which one is wrong. It is whether they are measuring differently or whether the models actually see you differently. Those two situations call for opposite responses, and telling them apart costs almost nothing.


FAQ

Why do AI visibility tools give different results for the same brand?
Most of the difference comes from methodology. Tools vary in how many queries they run, how they phrase them, how often they test, which systems they sample, and whether they count any brand mention or only sourced citations. The IAB identified more than 20 vendors producing materially different results for the same brand.

Is it normal for ChatGPT and Perplexity to say different things about my brand?
Yes. These systems retrieve from different source sets, weight recency differently, and require different levels of corroboration before recommending a brand. Consistent differences between them usually indicate that your evidence is concentrated in source types one system indexes and another does not.

How can I tell if a disagreement is a tool problem or a real one?
Run one important buying question by hand across ChatGPT, Claude, Gemini, and Perplexity in a single sitting, then repeat it the next day. If your manual runs agree but your vendors disagree, the gap is methodology. If your manual runs disagree the same way, the models genuinely see you differently.

Should I average my AI visibility scores across platforms?
Not without showing the spread alongside it. A brand scoring 60 on every engine and a brand averaging 60 from scores of 95, 80, 45, and 20 face completely different problems. The averaged number hides which engines represent gaps and which represent strengths.

What does it mean if only one AI system recommends my brand?
It usually indicates source concentration rather than a content deficiency. The engine recommending you is likely reaching a source type the others do not index or trust heavily. The fix is broadening the types of independent sources describing your brand, not publishing more on your own site.

Should I wait for AI visibility measurement standards to mature before starting?
No. The August 2026 IAB guidance is a floor rather than a finish line, and teams that start now will have a year of trend data before the vocabulary fully settles. Trend data over time is what converts a one-time snapshot into an understanding of your actual position.


Start here: axissuite.ai

Axis Suite by TrendAxis