AI Visibility: Why 4 Engines Score One Brand 10 to 100

June 23, 2026

Same brand, same skincare query, same afternoon: ChatGPT scored Tatcha a 10 out of 100 and Claude scored it a perfect 100. Not a rounding gap - the opposite answer, from two engines your customers treat as interchangeable.

I did not expect that. I ran our own AI-visibility audit - 31 direct-to-consumer brands, three categories, four engines (ChatGPT, Claude, Gemini, Perplexity) - mostly to stress-test our own scoring, expecting the engines to broadly agree on who the strong brands are. They did not. Not even close.

From what I have seen across our audit data, the gap between engines is not noise around some true score. Check a single engine and you have measured a quarter of your AI visibility, not all of it. Pick a different one and the same brand goes from invisible to dominant.

The one rule, up front

Your AI visibility score is only as honest as the four engines under it. There is a ChatGPT score, a Claude score, a Gemini score and a Perplexity score, and they disagree wildly. Roll them into one number without showing that spread and you bury the only honest part of the measurement.

Every AI engine gives the same brand a different visibility score

Here is the mean visibility score per engine, across the 20 brands we ran against all four. Same brands, same category queries, same day. The only thing that changed was the engine:

Engine Mean visibility (0-100)
ChatGPT 16
Perplexity 24
Gemini 35
Claude 44

Claude hands out visibility at nearly three times the rate ChatGPT does. That is not a rounding gap. ChatGPT was the stingiest engine for 16 of 20 brands; Claude the most generous for 14 of 20. They are not measuring the same thing - ChatGPT guards a short list, Claude recommends generously.

Now the number that actually matters. Take each brand's four engine scores and ask a simple question: do they land in the same rating tier - Low, Limited, Moderate, Strong, Excellent? Every single one of the 20 brands lands in a different tier depending on which engine you ask. Twenty out of twenty. There is not one brand in the set where the four engines agree on the tier.

Large companies keep announcing newer and larger gen AI models, but how do you know which model is truly the best if your evaluation methods aren't accurate or well studied?

Lingjia Tang Associate Professor of Computer Science and Engineering Tech Xplore, 2025
What an honest score needs

A single-engine number is not your AI visibility - it is one of four readings that disagree by whole tiers. A blended score is only worth trusting when it is built from all four engines and you can see the per-engine breakdown behind it. Hide that spread and the number is a coin flip wearing a lab coat.

One brand's AI visibility score swings from 10 to 100

Zoom into individual brands and it gets sharper. The average gap between a brand's best and worst engine in our set was 36 points. Here are the brands the engines fought hardest over:

Brand ChatGPT Claude Gemini Perplexity Spread
Tatcha 10 100 40 40 90
Four Sigmatic 10 70 50 40 60
Drunk Elephant 30 80 40 30 50
Adidas 30 70 40 30 40

Tatcha is the one that stops you. On ChatGPT it scores 10 - effectively invisible. On Claude it scores 100 - the only thing the model wants to talk about. Same brand, same skincare query, same afternoon. If Tatcha's team checked ChatGPT they would panic; if they checked Claude they would declare victory. Both would be wrong, because both measured a quarter of the picture.

This is the brand-recognition lottery we documented in the audit preprint, now showing up in our own first-party numbers. One engine is not a measurement. It is a draw - which is why one audit isn't enough.

Schema and AI crawler access did not explain AI visibility

Here is the part I did not expect, and the one that should change where you spend money. I checked whether the on-site signals everyone obsesses over - homepage structured data, AI-crawler access - explained any of the score. They did not.

  • Brands with homepage schema averaged 36. Brands with none averaged 35. A one-point difference, which is to say no difference.
  • Zero of the 31 brands blocked GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Crawler access was not the lever either - everyone already allows them.

So if it is not your schema and not your robots.txt, what is it? It is how densely your brand is referenced across the web these models trained and retrieved on. That is the whole thesis of our Mention Density Model: AI recognition is earned off your site, in the corpus, not on your homepage. Our own audit just failed to find any on-site signal strong enough to argue otherwise.

This is not permission to ignore schema. Clean structured data still fixes how engines describe you once they know you exist. But it will not get you into the answer. Being in the answer is a mention-density problem, and you solve it off your own domain.

The wider research lands in the same place. The Princeton GEO study (Aggarwal et al., 2023) tested what actually moves visibility inside generative engines and found the winners are content-credibility signals - citing sources, adding statistics, working in quotations - lifting visibility up to ~40% for the strongest. Notice what is not on that list: homepage schema.

Category and brand size shape AI visibility, not effort

Category matters more than effort. Mean visibility ran Sports & Fitness 47, Cosmetics & Beauty 35, Food & Nutrition 25. Functional food and supplements is simply a harder room to get recommended in - the engines hedge more on anything health-adjacent, and it shows up as a flat 20-point category penalty.

And real-world fame is not AI visibility. Lululemon and Nike saturate at 100. But Adidas sits at 43, Under Armour at 33, and serious DTC names - Vuori, Rhone, Fabletics - sit at 17, near-invisible despite real revenue and real customers. AI keeps a short, sticky list, and being big offline does not buy you a seat at it.

How to measure AI visibility across four AI engines

One audit, one engine, one snapshot is how almost everyone measures AI visibility right now. Our own data says that is close to worthless. Here is what I run instead, and you can do it for pennies:

  1. Score all four engines, never just one. ChatGPT, Claude, Gemini, Perplexity. A combined score is fine - the trap is trusting a single-engine reading, or a blend that hides the tier-level disagreement we found on every brand tested.
  2. Read your worst engine, not your best. A 100 on Claude and a 10 on ChatGPT is a ChatGPT problem, not a trophy. The floor is the work.
  3. Spend off-site, not on-site. Schema and crawler access explained none of the gap in our data. Mentions, reviews, third-party coverage and entity density did. Put the budget where the signal is.
  4. Pick category-aided queries you can actually win. Your category sets the ceiling. Find the rooms where mid-market brands get recommended and own those, instead of chasing "best [product]" terms the incumbents have locked for a decade.
  5. Re-run it. One snapshot is a mood, not a measurement. Engines drift between runs, so track the trend across several reads, never a single snapshot. The research backs this up: Schulte et al. (2026), "Don't Measure Once" argues AI-search visibility has to be measured repeatedly, and multi-prompt LLM evaluation research shows how sharply model results shift across prompt variations.

None of this is theory about how the models work. It is just what our own 31-brand audit made impossible to ignore: the engines do not agree, the disagreement is the signal, and the lever is off your website. Measure it that way and your visibility score finally means something, instead of swinging on whichever engine you happened to open.

Want a visibility score built on all four engines, with the per-engine breakdown behind it? Run GEOlikeaPro's Visibility checker - it scores ChatGPT, Claude, Gemini and Perplexity separately and rolls them into one number you can trust, because you can always see the spread underneath it.

FAQ

Do different AI engines give different brand visibility scores?

Yes, dramatically. In our 31-brand audit, mean visibility was 16 on ChatGPT versus 44 on Claude - nearly a 3x gap on the same brands, same queries, same day. And every one of the 20 brands we ran against all four engines landed in a different rating tier depending on which engine we asked. A single-engine score measures a quarter of the picture, and the four quarters disagree.

Should I average my AI visibility across engines into one score?

Yes, but only if it is built from all four engines and shown with the per-engine breakdown behind it. A blended score is fine; a single-engine number you treat as your AI visibility is not, and neither is an average that hides how far the engines disagree. Compute the score across ChatGPT, Claude, Gemini and Perplexity, then read your worst engine first - that is the one telling you where the work is.

Does structured data improve AI visibility?

Not for getting into the answer. In our audit, brands with homepage schema averaged 36 and brands with none averaged 35 - effectively no difference - and zero brands blocked AI crawlers, so access was not the lever either. Schema still matters for how accurately engines describe you once they know you exist, but presence in the answer is driven by off-site mention density, not on-page markup.

Which AI engine is hardest to get recommended by?

ChatGPT, by a wide margin. It scored brands a mean of 16 out of 100 and was the stingiest engine for 16 of the 20 brands we ran on all four. Claude was the opposite - a mean of 44 and the most generous engine for 14 of 20. If you only check ChatGPT you will understate your visibility; if you only check Claude you will overstate it.

Why does a big, well-known brand have low AI visibility?

Because real-world size is not AI visibility. In our data Lululemon and Nike hit 100, but Adidas sat at 43, Under Armour at 33, and large DTC names like Vuori, Rhone and Fabletics at just 17. AI answer engines keep a short, sticky list per category, and offline revenue does not buy a seat on it. Category also sets a hard ceiling: Sports & Fitness averaged 47, Beauty 35, and Food & Nutrition only 25.

Latest brands checked with our AI tools

The newest public reports from our AI visibility audit, hosting checker, and AI readiness checker - run yours to add your brand.

Stay ahead of AI search changes

Join store owners getting weekly GEO insights, AI search updates, and optimisation tips.

Get GEO tips →

Free tier · No credit card required