Cross-Provider Agreement on Discoverability Failures Isn’t the Same as Proof a Fix Transfers

A paper published in May 2026, with an arXiv identifier placing the manuscript at May 21 and v1 posted May 22, asked a question that matters directly for anyone managing AI visibility across more than one engine: when a brand has a discoverability problem, do different AI providers agree on where that problem sits, and does a fix that helps with one provider carry over to the others?

Start with the headline number, because it’s the one that gets quoted and the one most likely to be misread in either direction. Pooled agreement across providers on which stage a brand’s discoverability failure falls into is 95.1%, against a chance-agreement baseline of 93.0%. Read naively, that gap looks unremarkable, small enough to dismiss as noise. That reading isn’t quite right either. The high chance baseline exists because the pooled number is dominated by the most common, most basic discoverability failures, cases skewed enough that chance alone produces high agreement on them. The statistic built to correct for exactly this, Cohen’s kappa, comes in at 0.29 for this study, a modest, real result, not a sign the underlying observations are random.

The paper’s actual headline finding is more interesting than the raw percentage, and more useful than a blanket “it’s basically chance” correction of it. Agreement follows a clear prominence gradient: roughly 81% for well-known category leaders, rising to roughly 99.6% for obscure, long-tail regional brands. That pattern makes sense once you think through why it would happen. Providers overwhelmingly agree that an obscure, long-tail brand simply isn’t being surfaced, there’s very little room for disagreement when none of them are finding it in the first place. Well-known category leaders get more varied treatment, since there’s more nuanced territory, more finer-grained judgment calls, for providers to differ on.

The more granular number matters here too. Among the subset of 450 cases where at least one provider’s modal stage was Stage 2 or Stage 3, meaning that provider had cleared the most basic discoverability hurdle, agreement drops to 14.9%. Once a brand clears the first, most common failure mode, providers diverge substantially in how they assess what’s still wrong. That’s a meaningfully different, more nuanced picture than either the 95.1% headline or a flat “it’s close to random” dismissal would suggest.

Here’s the part worth being precise about, because it’s the part most likely to get oversold in either direction. This paper did not run a controlled intervention experiment. It measured recommendation sets, retrieval-attribution signals, and failure-stage agreement across providers. It did not apply a fix to one provider and measure the effect on another. The paper’s own discussion section goes further than its measured results: it infers that a shared, early-stage intervention would likely help long-tail brands across multiple providers, reasoning from how consistently those providers already agree on that failure stage. That’s a defensible, reasonable inference for the authors to draw. It is not the same thing as evidence that a transfer effect has actually been demonstrated.

This is a distinction worth holding onto deliberately, because conflating “providers agree on the diagnosis” with “a fix generalizes across providers” is exactly the kind of substitution that’s easy to make and expensive to get wrong. A team that assumes a fix transfers, based on diagnostic agreement alone, and skips separate verification on a second or third provider, could end up leaving a real, unaddressed problem in place for months, simply because the dashboard looked resolved everywhere once it was resolved on one engine. Nothing in a single-provider view would flag that the other engines were never actually checked.

This is also a useful way to think about what counts as a legitimate result versus an unfinished one. A team past the first discoverability stage, where cross-provider agreement drops to 14.9%, doesn’t have a failure on its hands just because the evidence doesn’t yet support a confident, single-cause explanation. Holding the question open, flagging it as something to keep watching rather than forcing a premature conclusion about why providers diverge, is a legitimate outcome in its own right. The alternative, manufacturing a tidy explanation the evidence doesn’t actually support, is the more common mistake, and it’s usually the one that looks more decisive in the moment and ages worse later.

There’s a genuinely useful, Axis-relevant takeaway buried in this paper that’s easy to miss if the only thing taken from it is the headline percentage: the evidence for cross-provider consistency is strongest at the most basic discoverability stage, which is exactly where the highest-leverage, lowest-cost fixes tend to live. That’s a reasonable basis for prioritizing basic discoverability work first. It’s not a reasonable basis for assuming any fix, once applied anywhere, is done everywhere. Treating each engine’s diagnosis as its own claim, and verifying a fix separately on each one past that first stage, is the direction worth building toward, distinct from claiming the question is already settled.

FAQ

What did the May 2026 paper actually measure?
It measured how consistently different AI providers agree on which stage a brand’s discoverability failure falls into, using recommendation sets and retrieval-attribution signals. It did not run a controlled experiment testing whether a fix applied to one provider transfers to another.

Is the 95.1% pooled agreement figure meaningful on its own?
Not fully on its own. It needs to be read against the paper’s 93.0% chance-agreement baseline, which is high because the pooled number is dominated by the most common, most basic discoverability failures. Cohen’s kappa, at 0.29, is the more precise measure of real agreement beyond chance.

What is the paper’s actual headline finding?
A prominence gradient: cross-provider agreement runs about 81% for well-known category leaders and climbs to about 99.6% for obscure, long-tail regional brands, since providers overwhelmingly agree an unfindable brand isn’t being found.

What does the 14.9% figure specifically represent?
It’s the same-stage agreement among the subset of 450 cases where at least one provider’s modal stage was Stage 2 or Stage 3, meaning that provider had cleared the most basic discoverability failure. Agreement is much lower among this subset than in the pooled figure.

Does this paper prove a fix that helps one AI provider also helps others?
No. The paper doesn’t test an actual intervention across providers. Its discussion section infers that an early-stage fix would likely help long-tail brands across providers, based on how consistently they agree on that failure stage, but that inference hasn’t been experimentally demonstrated.

What’s the practical takeaway for a business operating across multiple AI engines?
Basic discoverability fixes are worth prioritizing first, since that’s where cross-provider agreement is strongest. Anything beyond that first stage still needs separate verification on each provider rather than being assumed to transfer.

About Axis Suite
Axis Suite is the independent intelligence layer that tracks what AI says about your brand and the observable evidence behind it, retrieval, narrative, recommendation, and corroboration signals, across engines including ChatGPT, Claude, Gemini, and Perplexity. Built by TrendAxis, Axis Suite follows a brand’s path from Mentioned to Cited to Recommended to Chosen. Axis Suite is measurement infrastructure, not a marketing quick fix, every score is built to be explainable and defensible. Explore the Proof Center.