The Shortlist Index

A measurement method for answer-engine visibility that you can audit, reproduce and argue with. Published in full, because a metric nobody can check is not a metric.

Last updated · ~9 min read

Direct answer

How do you measure whether AI models recommend your brand? Build a fixed panel of 120–200 buyer-intent prompts, agreed in advance and published with the results. Run each prompt at least five times per answer engine so you can report a mean with a variance band instead of a single number. Stamp every data point with the model version that produced it. Then report three separate metrics — Recommendation Rate, Citation Share and Narrative Position — because visibility without favourable framing still loses deals.

The hidden denominator problem

Almost every AI visibility product on the market reports some version of share of voice: the percentage of AI answers in which your brand appears. It looks like the SEO metrics people already trust. It is not comparable to them, and the difference matters.

In traditional search, the denominator is knowable. There is a finite set of keywords with measurable volume, and you can say honestly that you rank for 340 of the 1,200 that matter. In answer engines there is no such set. The universe of possible prompts is effectively infinite. Every vendor therefore samples an arbitrary subset, and then reports the result as a percentage — which implies a denominator that was never disclosed and could not be justified if it were.

This is not a pedantic objection. It has already produced a visible failure. When a major provider shipped a significant model update in late 2025, citation counts moved across every dashboard in the industry at once, for reasons unrelated to any individual brand's relevance. Agencies spent that quarter explaining declines they had not caused and could not have predicted, using metrics that gave them no way to distinguish a platform event from a performance problem.

The fix is not a better score. It is a disclosed method. Every rule below is a constraint we accept in order to make the number auditable — and each one is also, not coincidentally, a question you can use to test any other vendor.

Rule 1 — Publish the denominator

We build a fixed panel of 120–200 prompts before any measurement happens. It is written jointly with you, because the language your buyers actually use is almost never the language in your marketing copy, and signed off in writing.

That panel is printed inside every report we produce. We never quote a percentage without showing the prompt set that generated it. If the panel changes — because your category shifts or a new competitor appears — the change is versioned and dated, and we re-baseline rather than quietly comparing across different denominators.

Panel composition is roughly: 40% category shortlist prompts, 25% head-to-head comparisons, 20% use-case and job-to-be-done phrasings, 15% alternative-seeking prompts.

Rule 2 — Buyer-intent prompts only

The panel contains no brand-name queries. Asking a model "what is Acme Analytics?" and celebrating that it answers correctly proves nothing — you handed it the answer inside the question.

What goes in the panel is what a buyer types before they know you exist:

  • Shortlist prompts — "best product analytics for a mobile-first team"
  • Head-to-head — "Mixpanel vs Amplitude for a 40-person startup"
  • Alternative-seeking — "cheaper alternatives to Amplitude"
  • Job-to-be-done — "how do I track feature adoption without a data team"
  • Constraint-led — "product analytics that is GDPR-safe and self-hostable"

This is the single biggest reason reported scores differ between vendors. A panel loaded with brand queries produces a flattering number that has no relationship to whether you win deals.

Rule 3 — Report a range, never a point

Large language models are non-deterministic. Ask the same question twice and you can get a different set of named vendors, in a different order, with different sources. A measurement taken once is a sample of size one presented as a fact.

We run every prompt a minimum of five times per engine per cycle — more where variance is high — and report the mean with its band. A brand at "34% ±11" is in a materially different position from one at "34% ±2", and a method that hides that difference is hiding the most decision-relevant thing in the dataset.

The practical consequence: we will sometimes tell you a month-on-month change is not significant. Vendors reporting clean single numbers cannot tell you that, because they have no basis on which to know.

Rule 4 — Stamp every data point with the model version

Each result we record carries the engine, the model version and the timestamp. When a provider ships an update, the discontinuity appears in your report as a labelled model event, annotated and separated from your own trend line.

We set this expectation in the proposal, not in an apology six months later. Scores will drop for reasons neither of us controls. The relationship survives that if — and only if — the reporting was built to show it honestly from the first cycle.

Rule 5 — Three metrics, because one is not enough

A single visibility score compresses three genuinely different things into one number, and the compression loses exactly the information you need.

The three reported metrics
MetricQuestion it answersHow it is calculatedTypical speed to move
Recommendation Rate When a buyer asks for a shortlist, how often are you on it? Share of shortlist-class prompts in which the brand is explicitly named, averaged across engines and runs 2–3 cycles
Citation Share When the model links a source, how often is it yours? Share of all outbound citations across the panel that resolve to the brand's own domain 4–6 cycles
Narrative Position When you are named, what are you named as? Classified framing of each mention — leader, challenger, budget option, legacy, niche — with sentiment 3–6 cycles

Narrative Position is the one almost nobody measures, and it is frequently the one costing money. Being named in every answer as "the cheaper option that lacks enterprise controls" is not a win. High visibility with poor framing actively damages pipeline, and you cannot see that in a share-of-voice percentage.

How this compares to the standard approach

Typical AI visibility reporting vs the Shortlist Index
DimensionTypical vendor reportingThe Shortlist Index
Prompt setUndisclosed, vendor-selectedPublished in every report, client-approved, versioned
Brand-name promptsUsually included, inflating the scoreExcluded entirely
SamplingOften a single run per promptMinimum five runs per engine per cycle
UncertaintyNot reportedVariance band on every figure
Model updatesAppear as unexplained score movementLabelled model events, separated from trend
MetricsOne composite visibility scoreThree metrics reported separately
Framing qualityNot measuredNarrative Position, classified and tracked

What actually moves the number

Measurement gets us hired; it is not the work. These are the levers, in rough order of impact.

  1. Off-site footprint before on-site anything. Models cite brands that other trusted sources already discuss — encyclopaedic references, large industry publishers, mainstream press, developer communities, code repositories. Earning presence in third-party content beats any amount of homepage rewriting.
  2. Placement in the comparison articles models quote. "Best X for Y" and "A vs B" formats are cited disproportionately because they are trivially extractable. Getting into the ones that already rank is the highest-leverage single tactic available.
  3. Entity clarity. Make it unambiguous what you are, who you serve and what you compete with. Consistent descriptions across the web, correct structured data, and your own honest comparison pages.
  4. Freshness cadence. Recently updated content attracts materially more citations than stale material. A quarterly refresh on the twenty pages that matter is cheap and compounds.
  5. Disclosed community presence. Community discussion carries real weight in retrieval. We participate with affiliation stated every time. We do not astroturf, and we will end an engagement rather than start.

What this method cannot do

Stated plainly, because a methodology page that only lists strengths is marketing.

  • It cannot attribute revenue. There is no reliable path today from an AI answer to a closed deal. We measure Recommendation Rate and say so; anyone promising traceable revenue is overselling.
  • It cannot cover the whole prompt space. A 200-prompt panel is a sample. It is a disclosed, defensible, stable sample, which is a different thing from a complete one.
  • It cannot predict model updates. We can label them after the fact and control for them going forward. We cannot forecast them.
  • It cannot make a weak product recommendable. Models increasingly reflect genuine sentiment. If the underlying reviews are poor, visibility work will surface that faster, not hide it.

The test we invite. Ask any vendor you are evaluating for their prompt panel and their variance data. Those two questions separate measurement from decoration, and you do not have to take our word for anything to ask them.

Next step

See the method run on your brand

We will run a real buyer prompt for your category and send you the verbatim answer, free. No call required.

Get a free finding →