The AI Recommendability Study (August 2026)
Our Brand Iceberg Model has a whole facet devoted to how brands earn their place inside AI-mediated buying journeys. The theory was grounded in interviews with senior marketers and a review of the available research. What it did not have was proof.
Meanwhile the market filled up with tools measuring visibility, discoverability and mentions. All useful. All a step or two away from what marketers actually need to know:
Does AI recommend me, and if not, what do I do about it?
So we built a study to test whether the below-the-waterline elements of the model, the parts built for AI rather than for people, actually drive whether a brand gets recommended.
What we did
| Categories | 7 (auto insurance, CRM software, cruise lines, retail banking in Australia, investing and brokerages, laundry detergent, accounting services) |
| Brands | 56, eight per category |
| AI responses logged | 2,612 |
| Platforms | ChatGPT, Gemini, Perplexity |
| Queries per category | 94 to 126 |
| Collection window | August 2026 |
| Scoring | Every response read and scored by hand |
| Brand audit | Seven elements, 0 to 4, public sources only, same standard for every brand |
| Reproducibility check | 210 query instances re-run three times, 630 responses |
Seven categories across consumer and B2B: auto insurance, CRM software, cruise lines, retail banking in Australia, investing and brokerages, laundry detergent, and accounting services.
Eight brands in each, from category leaders to niche upstarts, selected to span a range of maturity rather than to rank the category.
2,612 AI responses logged verbatim across ChatGPT, Gemini and Perplexity, from 94 to 126 distinct queries per category, collected in August 2026. Query design follows the five-type structure and intent mapping in the IAB's Measuring Visibility in the AI Era (August 2026).
Alongside the queries, we audited every brand's own digital estate against the below-the-waterline elements, scored 0 to 4 entirely from public sources. No client access, no privileged data, the same standard applied to every brand.
Every response was read and scored by hand against the Pick Rate scale. A stratified 10% sample of queries was run three times to measure reproducibility.
What we learned
Two things get measured, and they are not the same
The list. Do you show up, come to mind unprompted, make the shortlist?
The pick. Are you the single "this is the one for you" recommendation?
Visibility, unprompted mention, and shortlist presence turn out to be so closely related that they behave as one measure. Being the pick behaves as something separate. Half the brands we assessed were never the sole recommendation once, and plenty of them were on nearly every shortlist in their category. (All Pick Rate figures use the customer-scenario base; see Method notes)
| Within a category, how closely does it track plain visibility? | Correlation | Shared variance |
|---|---|---|
| Comes up unprompted | 0.97 | 94% |
| Makes the shortlist | 0.79 | 62% |
| Is the pick | 0.63 | 40% |
Three of these are close to the same measurement. The fourth is not. Pooled across all seven categories, the relationship between visibility and being the pick falls further, to 0.40, because categories differ in how winnable the pick is at all.
And two elements do most of the work, in sequence
Third-party authority is the gate. It determines whether AI treats a brand as eligible to be the answer at all.
Occasion-specific content is the conversion. Once eligible, it determines how often the brand actually wins.
Neither substitutes for the other. A brand with excellent content and thin outside validation rarely gets considered. A brand with strong validation and generic content gets shortlisted constantly and converts rarely.
| Combined score on third-party authority and occasion content | Chance of ever being the single recommendation |
|---|---|
| 1 | 9% |
| 2 | 25% |
| 3 | 51% |
| 4 | 77% |
Directional. The shape of the effect is solid; the exact figures carry real uncertainty at this sample size, and even the pessimistic end of the interval is a wide gap.
Comparison content, where a brand puts itself against named alternatives on its own site, shows a real lift. Fewer than a third of the brands publish any at all. The effect is more provisional than the first two and we are treating it that way.
There’s a third lever, largely under utilized
We audited seven elements. Two held up.
| Element audited | Effect on being the pick, in cold single-turn queries |
|---|---|
| Third-party authority | Predicts. Strongest at getting on the list. |
| Occasion-specific content | Predicts. Strongest at being the pick. Largest effect in the study. |
| Comparison content | Predicts, provisionally. Significant on one outcome, and we are treating it as directional. |
| Query ownership | No measured effect. |
| Content freshness | No measured effect. |
| FAQ content | No measured effect. |
| Identity architecture | Not testable here. Its job is preventing misidentification, and no outcome in this study measures identity resolution. |
These are nulls on the recommendation moment, not verdicts on the elements. Anything that works later in a conversation, once a buyer has narrowed to two options and started probing claims, would not surface in a study built on cold single questions.
Proof only counts if it travels
Not every award clears the gate. A verdict behind a paywall, or buried in a report that costs thousands of dollars to read, never reaches the systems doing the recommending.
In one category, two banks both scored well on third-party authority and both promoted their recognition hard. One converted into recommendations. The other never did, not once across 381 responses. The bank that lost had the better verdict, ranked first overall by a respected analyst firm. That report cost over a thousand dollars to read and never appeared. The winner's recognition came from a free comparison site that everyone cites. It appeared again and again.
What matters is not how prestigious the source is. It is whether the judgment gets repeated somewhere retrievable.
Content that is ownable, with fewer competitors claiming the same space, does substantially more work than category-generic content. If everyone makes the same claims and covers the same occasions, everyone gets credit and nobody gets picked.
That kind of content does not appear by accident. It comes from deciding which occasions a brand can actually own, which is a positioning decision before it is a content decision.
Distinctive content is worth roughly twice generic content
The same low score hides different problems
Brands AI cannot reliably identify. Brands that are simply absent. And brands that make every shortlist and are never chosen.
Most measurement treats these identically. They have different causes and the fixes are not interchangeable. Working on the wrong one is the most expensive mistake available here.
What this study does not cover
Every query was a single, cold question with no follow-up. Real buyers ask again, narrow and probe. Anything that works later in a conversation is invisible here, and we expect some things do.
AI recommendations are noisy. Asked the same question twice under identical conditions, the top brand changes roughly a quarter of the time. That makes our findings conservative rather than inflated, since noise pushes measured effects toward zero, but single-brand scores should be read as ranges rather than precise figures.
Eight brands per category is enough to find strong effects and not enough to find subtle ones. Six of the seven categories are US markets. Findings for third-party authority and occasion-specific content held as each category was added. Other results are directional.
This is observational. We measured brands at one point in time and found that higher scores go with higher Pick Rates. We have not yet observed a brand change its below-the-waterline position and watched its Pick Rate move in response. That study is worth running and we intend to run it.
Method notes
Platforms and models. ChatGPT (gpt-5.6-terra), Gemini (gemini-3.6-flash) and Perplexity (sonar), run through APIs with no account history, no personalization, and no conversation memory carried between queries. Perplexity was run on the base retrieval tier throughout rather than the deeper tier, held constant across all categories.
Collection window. August 2026.
Scoring. Every response read and scored by an AI assessment first, then checked by hand against the Pick Rate scale. Whole numbers only, no reconstructed entries.
Pick Rate base. All Pick Rate figures on this page are computed on the customer-scenario base: the eligible response set is every response to a customer-scenario query in that category, across all three platforms. This is the base recommended in the Pick Rate specification. Figures computed on a broader base are not comparable and would be roughly four times lower.
Reproducibility. A stratified 10% sample of query instances, weighted toward category and customer-scenario queries, was run three times. 210 query instances, 630 responses.
Below-the-waterline audit. Seven elements scored 0 to 4 from public sources only, using a fixed set of evidence standards applied identically across all 56 brands.