Pick Rate Specification

Version 1.0 | Mixtape Partners | 26 August 2026 | mixtapepartners.com/pick-rate

Pick Rate

A specification for measuring whether AI systems recommend a brand, rather than merely mention it.

USAGE NOTE: Free to use, implement, and cite. Attribution requested, permission not required.

What Pick Rate measures

Most AI visibility measurement counts appearances. A brand is mentioned, or it is not. A brand shows up on a list, or it does not.

That is worth knowing, and roughly a dozen tools now report it well. But it answers a different question than the one most marketers are actually asking. Appearing on a list of eight options is not the same commercial event as being told “this is the one for you.”

Pick Rate measures the second thing. It is the share of AI responses in which a brand is named as the single recommendation.

It is designed to sit alongside whatever presence metric you already use, not to replace it. Visibility, share of voice, mention rate, presence: keep the term and the measure you have. Pick Rate is the second axis, not a competing vocabulary for the first.

Definition

Pick Rate = the number of responses in which a brand is scored 4, meaning it is named as the single recommendation, divided by the number of responses in which that brand could have been named.

Expressed as a percentage, at the brand level, within a defined category.

Pick Rate is not meaningful across categories. Category structure determines how often any single answer is offered at all. Some categories resolve to one recommendation naturally; others resolve to shortlists because that reflects how the purchase genuinely works. Comparing a brand’s Pick Rate in one category against a brand’s Pick Rate in another measures the difference between the categories, not the difference between the brands. Always report Pick Rate against a category reference point.

Pick Rate was developed and validated across seven categories and 2,612 AI responses. See the study.

The denominator

Recommended definition: customer-scenario queries only. The eligible response set for a brand is every response to a customer-scenario query in that category, across every platform tested. A customer-scenario query describes a buyer's situation without naming a brand, and asks what to do about it.

This is the recommended base because it is where recommendation actually happens. A buyer describing a real need is asking to be matched. Category-level questions often resolve to explanation rather than a pick, and questions about a concept or feature frequently produce no brand recommendation at all. Including them measures the absence of a question rather than the failure of a brand.

A broader base is permitted, with a warning. Some implementations may prefer to score across the full eligible set: all category, customer-scenario and competitive queries, plus brand-anchored product queries and that brand's own direct-brand queries, excluding other brands' direct-brand queries. This is defensible, and it produces a more conservative figure.

It also produces a very different number. In our own data, moving from the customer-scenario base to the full eligible set reduces every brand's Pick Rate by roughly a factor of four, with the highest observed rate falling from 46% to 11%. Rank ordering is largely preserved; absolute values are not.

Pick Rates computed on different bases are not comparable. State which base you used, every time.

One exclusion applies to any base you choose: never include other brands' direct-brand queries in a brand's denominator. Doing so produces artifacts, including the counterintuitive result of a brand's overall presence scoring below its unprompted presence.

Scoring

Every response is read in full and scored once per brand present.

Pick Rate scoring scale
Score Meaning
4 Named as the recommendation. Only one brand per response can score 4. This is the score Pick Rate counts.
2 Genuinely recommended with real reasoning, but presented as one of several options.
1 Mentioned in passing, with little or no endorsement.
0 Not mentioned, or named as something to avoid.

Whole numbers only. No decimals, no averaging within a response.

There is no 3. This scale is ours. The IAB defines Recommendation Strength as a construct but deliberately does not prescribe a scale, noting that the boundary between active endorsement and passing mention is a methodological judgment call and requiring that providers disclose the rubric they use. This is ours, disclosed.

The gap between 2 and 4 is intentional. Being the answer is not one notch better than being an option, it is a categorically different outcome. A continuous scale invites scorers to split the difference on responses that are genuinely ambiguous, which is precisely where the one-per-response rule needs to hold. The gap forces the call.

The IAB's own distinction maps directly onto that boundary: a qualified recommendation ("a good option if you're on a budget") scores 2, an unconditional one ("the best option") scores 4.The one-per-response rule

At most one brand in any given response can score 4. This is the constraint that makes the metric mean something, and it is the rule almost no measurement tool applies.

A response earns a 4 for a brand when it uses explicit best-overall language, or places the brand in a clear first position with emphasis that separates it from the rest of the list.

Where a response lists several brands with no clear winner, no brand scores 4. A flat list is a shortlist. Treating the first item in an unranked list as a recommendation is the single most common way this measure gets inflated.

Sub-case qualifiers do not earn a 4. “Best for small teams” or “best if you already use Microsoft products” is a segmented recommendation within a shortlist, not a category answer. Score it 2.

When it is genuinely close, require all three signals. A response earns a 4 only when the brand is placed first, given the most developed treatment of any brand in the response, and referenced again in any closing summary or verdict. Where only one or two of those hold, score 2.

This is the boundary where scorers disagree most, and a stated rule matters more than the particular rule chosen. If you adopt a different threshold, publish it.

Query set design

Pick Rate is only as good as the queries behind it. A query set weighted toward direct-brand questions will overstate; a set weighted toward pure category questions will understate.

Use five query types, distinguished by what the buyer already knows when they ask.

The five query types
Type What the buyer already knows Example
Category The category and nothing else “What’s the best CRM software?”
Customer scenario Their situation, not the answer. This is the recommended base for Pick Rate. “Best car insurance for a family with teen drivers”
Direct brand A specific brand they are validating “Is GEICO any good?”
Product or offering A concept or feature, not a vendor “What is accident forgiveness and is it worth adding?”
Competitive A narrowed comparison “How does Carnival compare to other major cruise lines?”

Recommended minimum volume: 90 to 125 distinct query instances per category, covering all five types, with the brand roster held constant across every query where a brand could be named. These map onto the intent classes in the IAB’s Measuring Visibility in the AI Era: informational, recommendation, comparison and transactional. Report your mapping.

Platforms and versions

Report the specific model strings and the collection window. Not “ChatGPT, Gemini and Perplexity,” but the model versions and dates.

This matters more here than in most measurement work. The systems change underneath you without notice, and a Pick Rate figure with no version attached ages into a number nobody can interpret or reproduce.

Where a platform offers retrieval tiers, state which tier was used and hold it constant. Deeper retrieval tiers behave differently from single-pass tiers, and mixing them within a study makes the results uninterpretable.

Where possible, run through APIs rather than consumer interfaces, with no account history, no personalization, and no conversation memory carried between queries.

Variance and the replicate requirement

This is the part most implementations skip, and it is the part that determines whether a Pick Rate figure means anything.

AI recommendations are noisy. In our own testing across seven categories, asking an identical question twice under identical conditions changes which brand is named as the recommendation in roughly a quarter of pairs.

Noise of that size has two consequences.

At the brand level, aggregated across 90 or more queries, Pick Rate is considerably more stable than the per-query flip rate suggests. Rates are usable. Individual query results are not.

At the change-over-time level, the noise is larger than most of the movement anyone is trying to detect. A brand measured at 14% and remeasured at 19% has not necessarily improved.

So any implementation reporting Pick Rate should:

  1. Report the sample size behind every figure. Eligible responses per brand, not total responses collected across the category. This is the single most important disclosure and almost nobody makes it.

  2. Report ranges, not points, for any single brand’s Pick Rate.

  3. Do not claim a change smaller than the sample can support, and show the arithmetic.

  4. Run replicates where you can. Repeating a stratified sample of queries three times gives you a direct read on how noisy your own collection is. We use a 10% sample as a standing production rate.

  5. Double-score a sample and report the agreement rate. Pick Rate depends on judgment, which is the reason it cannot be fully automated and the reason it must be checked. Have a second scorer independently score a sample of responses without seeing the first scorer's results, and report how often the two agree on which brand, if any, scores a 4. An implementation that reports Pick Rate without an agreement rate is asking to be taken on trust.

Our own agreement rate. Across two categories, every response scored 4 was independently re-read and re-scored against the rubric by a second scorer working from the raw responses. 30 of 33 held. The three that did not were all the same error in the same direction: a flat list of options where the first-named brand had been treated as the recommendation, which the one-per-response rule exists to prevent.

This is a check on the 4s rather than full independent double-scoring of the whole corpus, and we report it as such. Agreement on the boundary that defines the metric is the number that matters most, and it is the one we measured.

What your sample size buys you

Scroll the table sideways to see all sample sizes.

Smallest change in Pick Rate you can detect, by eligible responses per brand
Expected Pick Rate n=36 n=90 n=180 n=360 n=720
5% 14 pts9 pts6 pts5 pts3 pts
10% 20 pts13 pts9 pts6 pts4 pts
15% 24 pts15 pts11 pts8 pts5 pts
25% 29 pts18 pts13 pts9 pts6 pts

Shaded cells are where a change of 10 points or less can be detected.

These figures are floors, calculated at 80% power. Run-to-run variance makes the real threshold somewhat larger. Read this before designing a tracking program: detecting a five-point change on a brand sitting around 5% takes roughly 300 eligible responses per brand. Most published AI visibility figures rest on far less than that. A Pick Rate published without a stated sample size is a number with unknown error attached.

What Pick Rate does not measure

Pick Rate is a single-turn measure. It captures the cold question, asked once, with no follow-up. Real buyers narrow, probe, and ask again. Anything that only takes effect in the third turn of a conversation is invisible to this metric, and some things almost certainly do.

Pick Rate is a snapshot of recommendation, not of outcome. It does not measure traffic, consideration in the buyer’s own head, or purchase.

Pick Rate says nothing about why a brand is or is not the pick. Diagnosis is a separate exercise.

Reporting standard

A Pick Rate figure should always be published with:

  • The category and the brand roster

  • The query set size and the distribution across the five types

  • Platform model versions and the collection window

  • Eligible responses per brand

  • A category reference point: median, or the range across the roster

Citation

Pick Rate (v1.0). Mixtape Partners, 2026. mixtapepartners.com/pick-rate

We developed this metric and we are giving it away. We would rather it became the standard way the industry talks about AI recommendation than remained a term only we use. Implement it, extend it, argue with it. If you improve on it, tell us.