I Expected Level 4 AI Visibility. I Landed at Level 2. Here Is Why That Gap Changed Everything.

The first time I tested my own brand across all four major AI platforms, I was confident about where I stood.

I had published content consistently. My positioning was clear. My product was genuinely differentiated. I had a Capterra listing, a G2 profile, and active engagement across multiple communities. By any reasonable assessment, AI should have known about me and recommended me with confidence.

I expected Level 4. Trusted. Recommended with independent evidence corroborating the claims.

I ran the diagnostic. Same buyer-intent queries across ChatGPT, Claude, Gemini, and Perplexity. Checked all four platforms. Added constraints. Checked descriptions. Looked for independent evidence.

I landed at Level 2. Intermittent. Appeared on some platforms. Absent on others. Dropped out the moment queries got specific.

The gap between where I thought I was and where I actually was taught me more about AI visibility than anything else I have learned building in this space.

Why the Gap Exists

The gap happens because testing yourself on one platform with one broad prompt feels like visibility.

You ask ChatGPT about your category. Your brand appears. You feel validated. You move on to the next thing on your list.

But consistent AI visibility requires passing a much higher bar.

Do you appear on all four platforms, not just one? Do you survive when the buyer adds constraints like company size, industry, or use case? Is the description accurate across every platform, or does one describe you correctly while another references capabilities you retired? Does independent evidence corroborate the recommendation, or is AI just echoing your own website content?

Most brands have never asked those questions. They tested once, on one platform, with one broad prompt, and treated the result as their baseline.

That single data point is not visibility. It is a sample of one in a system that changes by platform, by query, and by day.

What Level 2 Actually Means

Level 2 is Intermittent. It means AI has some awareness of your brand but not enough recommendation confidence to include you reliably.

In practical terms, Level 2 looks like this: you appear on ChatGPT for a broad category query but you are absent from Gemini for the same query. You show up when a buyer asks “what are the best platforms for [your category]” but you disappear when they follow up with “which of those is best for a small team in [specific industry].” You are in the candidate pool but you do not persist through specificity.

The cause at Level 2 is usually not content quality. It is consistency. Category language varies across your website, directory listings, and third-party profiles. One source describes you as “AI marketing software.” Another describes you as “competitive intelligence.” A third describes you as “visibility analytics.” AI systems see three different positioning stories and cannot form a stable recommendation.

This was my exact problem. My content was strong. My positioning was clear on my own site. But the broader evidence landscape (directory listings, third-party mentions, schema markup, category alignment across sources) was not consistent enough for AI to form a durable belief about what category I belonged in.

What I Would Have Done Wrong

If I had trusted my initial assumption (Level 4), I would have invested in Level 4 fixes: building more independent evidence, pursuing analyst coverage, expanding review generation.

None of those investments would have moved the needle because they address the wrong layer. Level 4 fixes assume Level 2 and Level 3 are already working. They assume AI consistently includes you and describes you accurately. Those assumptions were false.

Building independent evidence for a brand that AI cannot consistently find or accurately describe is like advertising a restaurant that is not on any map. The investment goes into awareness for something people cannot locate.

The level determines the fix. The fix at Level 2 is consistency: align category language across every source AI draws from. Make sure your directory listings, schema markup, third-party profiles, and website all tell the same story about what you are and what category you belong in.

That is a fundamentally different investment than what Level 4 requires. And the only way to know which investment is right is to know which level you are actually on.

How I Tested and What Changed

The test itself is simple. It takes five minutes.

Ask ChatGPT, Claude, Gemini, and Perplexity about your category. Work up the ladder. Stop at the first level where the honest answer is “no.”

Do you appear at all? If no, Level 1. Retrieval is the issue.

Do you appear on all four platforms and survive follow-up questions? If no, Level 2. Recommendation consistency is the issue.

Is the description accurate and consistent across all platforms? If no, Level 3. Narrative alignment is the issue.

Does independent evidence corroborate the recommendation? If no, Level 4. Evidence is the issue.

Do you persist through model updates and competitive changes? If no, you have not yet reached Level 5.

After identifying my actual level, I shifted my entire approach. Instead of publishing more content (which is the default instinct when visibility feels low), I focused on consistency. I updated directory listings. I aligned category language across every source. I cleaned up schema markup. I made sure every reference to my brand told the same story.

Those fixes addressed the actual problem. Content was not the problem. Consistency was.

The Two-Level Gap Is Universal

I have since run this same diagnostic with other brands. The gap is remarkably consistent. Most brands overestimate their maturity level by approximately two levels.

A brand that thinks it is Trusted (Level 4) is usually Intermittent (Level 2). A brand that thinks it is Recognized (Level 3) is usually Invisible (Level 1). The overestimation follows the same pattern because the cause is the same: testing on one platform with one broad prompt creates a false baseline.

The 30% consistency statistic from AirOps research reinforces this. Only 30% of brands maintain consistent visibility across AI sessions. The other 70% flicker in and out depending on the prompt, the platform, and the day. Most of that 70% believe they are in the 30%.

The Uncomfortable Lesson

The hardest part of this experience was not the diagnostic. It was accepting the result.

I built Axis Suite to help brands understand their AI visibility. And my own brand was two levels below where I assumed it was. That is humbling.

But it was also the most useful thing I learned. Because once I knew the actual level, I knew the actual fix. And the fix was not what my instinct would have chosen.

The level you land on might be uncomfortable. The fix it points to will be more useful than anything the score alone could tell you.

Stop watching the number. Start locating yourself on the curve. The curve tells you where you are. The level tells you what to build next.


Frequently Asked Questions

Why do most brands overestimate their AI visibility maturity level?
Because testing on one platform with one broad prompt feels like visibility. Finding your brand in a single ChatGPT answer creates a false baseline. Consistent presence across all four major platforms, surviving follow-up questions, with accurate descriptions and independent evidence, is a much higher bar.

What is the difference between Level 2 and Level 4?
Level 2 (Intermittent) means AI knows about you but does not include you consistently across platforms or through specific queries. Level 4 (Trusted) means AI recommends you with confidence and independent sources corroborate the recommendation. The gap between them requires two layers of improvement: narrative consistency (Level 3) and independent evidence (Level 4).

How do I test my actual maturity level?
Ask ChatGPT, Claude, Gemini, and Perplexity about your category. Check whether you appear on all four, survive follow-up questions with constraints, are described accurately, and have independent evidence backing the recommendation. Where you drop out reveals your actual level.

What should I do if I land at a lower level than expected?
Focus on the fix for your actual level, not the level you expected. Level 1 requires crawlability fixes. Level 2 requires category language consistency. Level 3 requires narrative alignment across sources. Level 4 requires independent evidence building. Working on the wrong level wastes time and budget.

Can the maturity level change quickly?
Level 1 to Level 2 can happen in days if the fix is structural. Level 2 to Level 3 typically takes weeks as consistency propagates. Level 3 to Level 4 takes weeks to months as independent evidence accumulates. Level 4 to Level 5 is measured in quarters.

Is a two-level gap always the case?
It is the most common pattern but not universal. Some brands are accurately self-assessed. Some are three levels off. The gap tends to be larger for brands that have only tested on one platform with one type of query and smaller for brands that regularly test across multiple platforms with varied buyer-intent prompts.


Start here: axissuite.ai

Axis Suite by TrendAxis