The AI Recommendability Study (August 2026)

Our Brand Iceberg Model has a whole facet devoted to how brands earn their place inside AI-mediated buying journeys. The theory was grounded in interviews with senior marketers and a review of the available research. What it did not have was proof.

Meanwhile the market filled up with tools measuring visibility, discoverability and mentions. All useful. All a step or two away from what marketers actually need to know:

Does AI recommend me, and if not, what do I do about it?

So we built a study to test whether the below-the-waterline elements of the model, the parts built for AI rather than for people, actually drive whether a brand gets recommended.

What we did

Categories7 (auto insurance, CRM software, cruise lines, retail banking in Australia, investing and brokerages, laundry detergent, accounting services)
Brands56, eight per category
AI responses logged2,612
PlatformsChatGPT, Gemini, Perplexity
Queries per category94 to 126
Collection windowAugust 2026
ScoringEvery response read and scored by hand
Brand auditSeven elements, 0 to 4, public sources only, same standard for every brand
Reproducibility check210 query instances re-run three times, 630 responses

Seven categories across consumer and B2B: auto insurance, CRM software, cruise lines, retail banking in Australia, investing and brokerages, laundry detergent, and accounting services.

Eight brands in each, from category leaders to niche upstarts, selected to span a range of maturity rather than to rank the category.

2,612 AI responses logged verbatim across ChatGPT, Gemini and Perplexity, from 94 to 126 distinct queries per category, collected in August 2026. Query design follows the five-type structure and intent mapping in the IAB's Measuring Visibility in the AI Era (August 2026).

Alongside the queries, we audited every brand's own digital estate against the below-the-waterline elements, scored 0 to 4 entirely from public sources. No client access, no privileged data, the same standard applied to every brand.

Every response was read and scored by hand against the Pick Rate scale. A stratified 10% sample of queries was run three times to measure reproducibility.

What we learned

Two things get measured, and they are not the same

  1. The list. Do you show up, come to mind unprompted, make the shortlist?

  2. The pick. Are you the single "this is the one for you" recommendation?

Visibility, unprompted mention, and shortlist presence turn out to be so closely related that they behave as one measure. Being the pick behaves as something separate. Half the brands we assessed were never the sole recommendation once, and plenty of them were on nearly every shortlist in their category. (All Pick Rate figures use the customer-scenario base; see Method notes)

Within a category, how closely does it track plain visibility? Correlation Shared variance
Comes up unprompted0.9794%
Makes the shortlist0.7962%
Is the pick0.6340%

Three of these are close to the same measurement. The fourth is not. Pooled across all seven categories, the relationship between visibility and being the pick falls further, to 0.40, because categories differ in how winnable the pick is at all.

And two elements do most of the work, in sequence

Third-party authority is the gate. It determines whether AI treats a brand as eligible to be the answer at all.

Occasion-specific content is the conversion. Once eligible, it determines how often the brand actually wins.

Neither substitutes for the other. A brand with excellent content and thin outside validation rarely gets considered. A brand with strong validation and generic content gets shortlisted constantly and converts rarely.

Combined score on third-party authority and occasion content Chance of ever being the single recommendation
19%
225%
351%
477%

Directional. The shape of the effect is solid; the exact figures carry real uncertainty at this sample size, and even the pessimistic end of the interval is a wide gap.

Comparison content, where a brand puts itself against named alternatives on its own site, shows a real lift. Fewer than a third of the brands publish any at all. The effect is more provisional than the first two and we are treating it that way.

There’s a third lever, largely under utilized

We audited seven elements. Two held up.

Element audited Effect on being the pick, in cold single-turn queries
Third-party authorityPredicts. Strongest at getting on the list.
Occasion-specific contentPredicts. Strongest at being the pick. Largest effect in the study.
Comparison contentPredicts, provisionally. Significant on one outcome, and we are treating it as directional.
Query ownershipNo measured effect.
Content freshnessNo measured effect.
FAQ contentNo measured effect.
Identity architectureNot testable here. Its job is preventing misidentification, and no outcome in this study measures identity resolution.

These are nulls on the recommendation moment, not verdicts on the elements. Anything that works later in a conversation, once a buyer has narrowed to two options and started probing claims, would not surface in a study built on cold single questions.

Proof only counts if it travels

Not every award clears the gate. A verdict behind a paywall, or buried in a report that costs thousands of dollars to read, never reaches the systems doing the recommending.

In one category, two banks both scored well on third-party authority and both promoted their recognition hard. One converted into recommendations. The other never did, not once across 381 responses. The bank that lost had the better verdict, ranked first overall by a respected analyst firm. That report cost over a thousand dollars to read and never appeared. The winner's recognition came from a free comparison site that everyone cites. It appeared again and again.

What matters is not how prestigious the source is. It is whether the judgment gets repeated somewhere retrievable.

Content that is ownable, with fewer competitors claiming the same space, does substantially more work than category-generic content. If everyone makes the same claims and covers the same occasions, everyone gets credit and nobody gets picked.

That kind of content does not appear by accident. It comes from deciding which occasions a brand can actually own, which is a positioning decision before it is a content decision.

Distinctive content is worth roughly twice generic content

The same low score hides different problems

Brands AI cannot reliably identify. Brands that are simply absent. And brands that make every shortlist and are never chosen.

Most measurement treats these identically. They have different causes and the fixes are not interchangeable. Working on the wrong one is the most expensive mistake available here.

THE LIST AND THE PICK Where a brand sits tells you which problem it has THE PICK are you the single recommendation THE LIST do you show up, come to mind unprompted, make the shortlist RARELY SURFACES / OFTEN THE PICK Sharply positioned, under-reached When you do come up, you win. You just do not come up often enough. Focus has done its job. Reach has not. The move: third-party authority. Get judged somewhere that gets republished. ON EVERY LIST / OFTEN THE PICK Chosen Eligible and converting. Both jobs done, and the position most brands in this study never reached. The move: defend it. Watch for competitors mapping the occasions you own. RARELY SURFACES / RARELY THE PICK Absent Either AI cannot reliably identify you, or it can and you are simply not there. Two problems wearing one score. The move: settle identity first. Nothing downstream works until AI knows who you are. ON EVERY LIST / RARELY THE PICK On every list, never the pick Trusted, present, consistently described as credible. And not the answer. The move: occasion-specific content. If that is already strong, the problem is positioning. A two-axis map. The horizontal axis is the list: whether a brand shows up, comes to mind unprompted, and makes the shortlist. The vertical axis is the pick: whether it is named as the single recommendation. Four positions: sharply positioned but under-reached, chosen, absent, and on every list but never the pick. The list and the pick

What this study does not cover

Every query was a single, cold question with no follow-up. Real buyers ask again, narrow and probe. Anything that works later in a conversation is invisible here, and we expect some things do.

AI recommendations are noisy. Asked the same question twice under identical conditions, the top brand changes roughly a quarter of the time. That makes our findings conservative rather than inflated, since noise pushes measured effects toward zero, but single-brand scores should be read as ranges rather than precise figures.

Eight brands per category is enough to find strong effects and not enough to find subtle ones. Six of the seven categories are US markets. Findings for third-party authority and occasion-specific content held as each category was added. Other results are directional.

This is observational. We measured brands at one point in time and found that higher scores go with higher Pick Rates. We have not yet observed a brand change its below-the-waterline position and watched its Pick Rate move in response. That study is worth running and we intend to run it.

Method notes

Platforms and models. ChatGPT (gpt-5.6-terra), Gemini (gemini-3.6-flash) and Perplexity (sonar), run through APIs with no account history, no personalization, and no conversation memory carried between queries. Perplexity was run on the base retrieval tier throughout rather than the deeper tier, held constant across all categories.

Collection window. August 2026.

Scoring. Every response read and scored by an AI assessment first, then checked by hand against the Pick Rate scale. Whole numbers only, no reconstructed entries.

Pick Rate base. All Pick Rate figures on this page are computed on the customer-scenario base: the eligible response set is every response to a customer-scenario query in that category, across all three platforms. This is the base recommended in the Pick Rate specification. Figures computed on a broader base are not comparable and would be roughly four times lower.

Reproducibility. A stratified 10% sample of query instances, weighted toward category and customer-scenario queries, was run three times. 210 query instances, 630 responses.

Below-the-waterline audit. Seven elements scored 0 to 4 from public sources only, using a fixed set of evidence standards applied identically across all 56 brands.