Measuring whether AI answers propose your product is a sampling problem, not a lookup. A single answer is a real observation of one event, but it cannot estimate how often you appear, and the churn is large enough that a before-and-after pair settles nothing on its own. What works is ordinary and unglamorous: a fixed panel of prompts, run repeatedly under recorded conditions, with each outcome coded in its own column, reported per engine and read as a distribution. The published run-count guidance comes from a single study in a single market and is a floor rather than a settled standard. This page sets out the protocol. It covers shortlist standing specifically; general citation measurement is set out in measuring AI citations and is not re-derived here.
Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked platform documentation, quoted. Recommendation means a workflow, threshold or coding rule that CoreAEX prescribes and no platform documents - which is most of this page, by nature. Research findings are attributed in the prose with their sample, dates and the outcome they measured, and are not tagged. Untagged text is ordinary explanation.
What the Panel Measures
Code each outcome in its own column, and do not combine them into a score before the raw coding exists. The pillar separates ten outcomes that are routinely collapsed. A prompt panel directly codes six of them and adds two run-status fields needed for valid denominators: no generated answer and run void. Retrieval requires different evidence, while click, consideration-set inclusion, and selection or purchase are buyer-side outcomes the panel cannot observe.
| Column | What it records | Why it is separate |
|---|---|---|
| Mentioned | Your product name appears anywhere in the answer text | The broadest presence outcome, and the one most often mistaken for the others |
| Recommended | The answer presents your product as suitable for the stated need | A name can appear as a competitor's comparison point or a criticism |
| In the proposed set | Your product appears in the candidate list the answer offers for evaluation | This is the outcome the commercial question is actually about |
| Position | Ordinal placement within that list | Unstable across runs; useful only in aggregate, never from one answer |
| Cited | The answer links or attributes to a URL, and which domain | Citation and recommendation are distinct outcomes; either can occur without the other |
| Persisted | Still present after a defined follow-up turn | A different question from first-turn presence, and worth separating if you ask follow-ups at all |
| No generated answer | The engine returned links or declined rather than producing an answer | A documented product behaviour, not a failed run - see below |
| Run void | The run failed for a reason unrelated to the question: fetch error, login wall, timeout | Excluded from denominators, counted and reported separately |
The last two rows matter more than they look. Google documents that in some cases, AI Mode will provide a set of web links if there's not high enough confidence in the quality or helpfulness of an AI response
. Documented A run that produces no generated answer is a real observation of the product's behaviour on that prompt, and folding it in with genuine failures throws away a finding. Keep void runs separate too, and report the void rate - a panel with a quarter of its runs voided is telling you something about access before it tells you anything about visibility.
Building the Prompt Panel
Changing prompts breaks a clean like-for-like trend for those prompts. Write the prompts once and version the file. Preserve an unchanged anchor set that carries across rounds, and treat edited or added prompts as a new panel version, reported separately - a changed panel can still support analysis through the stable overlap, but it cannot be read as one continuous series. Recommendation
Four prompt types answer four different questions, and averaging across them hides the answer to each:
- Category prompts - "best contract management software." Broad, competitive, and the least informative about fit.
- Use-case prompts - "software for tracking supplier certifications." Closer to how buyers actually ask.
- Constraint-heavy prompts - "for a 500-person fintech in the EU that needs SSO and SOC 2." The type where published evidence about your product's fit does most of the work.
- Comparison prompts - "X versus Y for mid-market teams." A different retrieval job, and often a different source mix.
Keep the prompts unbranded where the question is discovery. Naming your product in the prompt measures whether the engine can describe something it has been handed, which is a different and much easier question than whether it proposes you unprompted. One preprint puts the gap starkly: Amit Prakash Sharma's "The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries" (arXiv, January 1, 2026; based on M.Tech thesis research at the Indian Institute of Technology Patna, per the arXiv listing) tested 112 Product Hunt-featured products across 2,240 queries to two large language models and reports a recognition rate - the product named in the prompt - of 99.4% on ChatGPT and 94.3% on Perplexity, against an organic discovery rate in unbranded queries of 3.32% and 8.29% respectively. The paper's own limitations narrow it further: the sample is drawn from Product Hunt and so skews toward developer tools, productivity and AI products, and only two models were tested, with the author noting that Claude, Gemini and others might show different patterns
. Those rates should not be read as rates for established software. The distinction they illustrate is the durable part: recognition and discovery are different measurements, and only one of them is about your shortlist standing.
How Many Runs, Over What Window
The published guidance is at least seven runs per prompt-engine cell per day for brand-detection monitoring, at least eight per prompt-engine cell per day where source coverage matters, aggregated over a rolling two-to-four-week window. The unit matters: seven runs spread across four engines is not what the guidance describes. And note what it was measured on. This evidence concerns brand detection and source coverage, not B2B shortlist inclusion; this protocol carries the schedule across as a starting floor, not as a validated shortlist-specific threshold.
Schulte, Bleeker and Kaufmann of the University of St. Gallen, in "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv preprint submitted April 8, 2026; not peer-reviewed), argue that visibility should be characterised as a distribution rather than a single-point outcome
. Their recommendation reads at least 7 runs per prompt per day for brand visibility monitoring, and at least 8 runs when source-level coverage matters
, and their design ran the repetitions to each engine - 8 prompts per campaign queried up to 10 times to all 4 engines in succession
- which is why the operational unit is the prompt-engine cell rather than the prompt alone. Their window comes from a separate temporal dataset tracked daily over 45 to 46 days, where standard error drops below 0.10 at ten days and below 0.05 at 24 days - hence rolling aggregation over two to four weeks is therefore recommended to obtain per-brand estimates that are both statistically stable and representative of sustained visibility rather than momentary snapshots
. The study covers four engines, four commercial verticals, eight prompts per campaign and a Swiss, German-language market, and the authors note results may not generalise to other regional or linguistic markets
.
A second source adopts that schedule but does not independently confirm the numbers, and the difference is worth keeping straight. AirPulse, an AI-visibility vendor, tells practitioners to use seven brand-detection runs and eight source-coverage runs per prompt-engine cell as a starting floor
and states plainly that the recommendation comes from the paper Don't Measure Once
. Its own production cohort - 30,504 response observations across four engines over a 30-day window - supports the need for repeated, time-distributed measurement, and it says its study supports the direction - repeat measurements - but does not independently prove that seven and eight are universally sufficient
. What it adds is the operational reason for spreading the runs: Running eight times in ten minutes measures short-term generation variability. It does not capture day-to-day retrieval, model or index changes.
So: one study behind the thresholds, one vendor operationalising them. Raise the counts where the prompt is commercially important, where a result sits near a decision threshold, where engines disagree, or where the set of cited sources is still moving between rounds. Recommendation
Where a Measurement Budget Actually Buys Reliability
There is a real design tension here, and the honest answer is that it has been measured once, on a different outcome. Dmitrij Żatuchin's variance-components decomposition (arXiv preprint submitted July 14, 2026; not peer-reviewed) analysed 12,933 responses across 20 Central and Eastern European brands, eight languages and three models, and embedded the result in a decision study - how many repeats, paraphrases, models and languages to buy for a target reliability. Ranked by the variance reduction each block buys, the paper finds languages first (a reduction of 0.00462), then models (0.00167), then paraphrases (0.00125), and repeats last (0.00030) - language diversity cuts relative-error variance roughly fifteen times as much as five additional repeats do. It reports brand-ranking reliability of Eρ² ≈ 0.01 for a single answer, rising to about 0.36 at the full crossed 8-language, 3-model, 15-paraphrase design, and concludes that reliability is bought by breadth across languages and models, not by depth of repetition
.
Scope that carefully before spending against it. Its outcome variable was per-response sentiment polarity toward a brand, not shortlist inclusion - and the authors report that 91.9% of responses score exactly neutral
, so the decomposition ran on a near-degenerate outcome. Two of its three models - GPT-5.2 and Gemini 3 Flash - were queried in parametric mode without retrieval; only the third, Perplexity, used grounded retrieval. The authors name refitting the model on a recommendation indicator as future work, and expect that the brand-object signal is expected to be larger on this outcome than on the near-degenerate sentiment outcome
. Its brands were Central and Eastern European across eight languages, a design built to make language variation visible - which is the axis its conclusion favours. It is an outcome-specific demonstration that a fixed query budget may buy more reliability through breadth than through additional repeats. Whether that allocation carries over to shortlist inclusion is untested.
Read alongside the run-count guidance, the practical reading is a sequence rather than a contradiction. Repetition is what makes a single prompt's result interpretable at all; breadth - more prompts, more engines, and more markets if you sell into them - is what makes the overall picture reliable. A panel that runs one prompt fifty times on one engine has bought precision about something very narrow. Recommendation
What to Record on Every Run
Conditions you did not record are conditions you cannot rule out later. Every one of the fields below is either a documented input to the answer or a known source of variation. Recommendation
| Field | Why it is on the list |
|---|---|
| Engine and product surface | The four engines disagree consistently; AI Mode and AI Overviews are different surfaces |
| UTC timestamp | The scope anchor for every figure you will later report |
| Prompt ID and panel version | Ties the run to a frozen panel; makes an accidental edit visible |
| Repetition number | Distinguishes within-day repeats from across-day observations |
| Geography and language | OpenAI documents that a VPN or network location may affect the approximate location inferred from your IP address, and that if memory is enabled, saved location information may also influence search resultsDocumented |
| Account state - logged in or out | Google documents Personal Intelligence referencing Search Services History; OpenAI documents that if memory is enabled, ChatGPT may use relevant saved memories when rewriting a search queryDocumented |
| Session freshness | A second question in one thread can be answered from the first answer rather than from a new retrieval |
| Browsing or grounding active | Changes what the answer is drawn from |
| Model label | Record "not exposed" where the interface shows none, rather than leaving it blank or guessing |
| Full response text and displayed sources | The raw artefact. Coding decisions get revisited; responses cannot be re-fetched |
Two of those need saying directly. Pick one account state and stay in it. A logged-in session with memory enabled and a fresh logged-out session are different measurement instruments; neither is more correct, and comparing across them silently is a common way to manufacture a change. And keep the raw response. Coding rules get refined mid-round more often than anyone plans for, and a stored response can be recoded while a summary cannot.
Coding a Response
Write the coding rules before the first run, and treat a mid-round change as a new baseline. Recommendation Four decisions cause most of the disagreement between coders:
- What counts as "the proposed set." Answers rarely present a tidy numbered list. Decide in advance whether a product named in a closing caveat, or in a paragraph about a different need, is inside the set or outside it.
- Mentioned versus recommended. A name appearing inside a competitor's comparison, or in a sentence about what a product does badly, is a mention and not a recommendation. This is the single most common coding slip, and it inflates in the flattering direction.
- Cited versus displayed. Record the domain of every source the interface shows, and separately whether the answer attributes a specific claim to it. A source panel is not the same as an attribution.
- Follow-ups. If you ask them, ask the same one every time and code it as its own column. An unplanned follow-up is a different prompt.
Open the sources the answer displayed. This is the step teams skip and the one that pays: it tells you which of your own properties are in play, whose third-party pages are standing in for you, and whether what the answer says about your product matches what the cited page actually says. A wrong fact about you is a different problem from an absence, and it needs a different fix.
What Your Own Data Can and Cannot Answer
Your analytics answer a different question from your panel, and neither substitutes for the other.
Google's generative AI performance report in Search Console shows data about how your site performs in generative AI features on Google Search
, covering AI Overviews and AI Mode, and lets you see how your organic impressions from Search generative AI features changes over time
. It groups those impressions by pages - the final URL linked by a generative AI feature after any redirects
- countries, dates and devices. All dates are in Pacific Time Zone (PT)
, and Google notes that the usual data limitations (1,000 row limitation, time period, etc) for the Search performance report also apply to this report
. It is on partial rollout: the report's own note states, We're rolling out this report to a subset of website owners, allowing for thorough testing before rolling it further
, and its troubleshooting section separately adds, Not all properties have access to the report, as we're rolling out over time
. Documented
That is genuinely useful, and it is a record of link impressions. It can show when links to your site received impressions in AI Overviews or AI Mode and group them by page, country, date or device. It does not expose the prompt, the answer wording, whether you were proposed as a candidate, or your recommendation status - and it covers Google's features only, not ChatGPT, Claude or Perplexity. The panel exists precisely because that view does not exist in anyone's analytics.
Server, CDN and edge logs answer a third question again - whether a given engine's fetcher reached your pages. A log entry shows that a request reached the logged layer; it does not show indexing, retrieval, content use, citation or recommendation. That layer is covered in the AI-crawler guide.
When a Difference Is Worth Reporting
Decide what would count as a change before you look at the result. Comparisons chosen after seeing the numbers should be labelled exploratory or post hoc - they are still comparisons, but churn of the size this cluster documents, combined with the number of cuts available, makes a chance finding easy to select. The scale of that churn, measured across four different time intervals, is set out in why AI vendor shortlists change across prompts, buyers and markets - this section assumes that evidence rather than re-deriving it. Recommendation
Five rules make a reported difference defensible.
Report per engine, never pooled. Engines disagree consistently enough that an average across them describes no system that exists.
Report a distribution, not a point. "Present in 22 of 28 runs across the window" says what happened. "We appear in ChatGPT" does not, and cannot be checked.
Hold the conditions fixed across the comparison - same panel version, same engines, same account state, same geography, same coding rules. If one of them changed, that is the finding until shown otherwise.
Separate presence from composition. Whether you appear and where you sit in the list behave differently; a stable presence rate alongside a moving position is an ordinary result, not a contradiction.
Say what a change is consistent with, not what caused it. Between two rounds the engines changed, your competitors changed, the web changed and you changed. A movement after a content change is consistent with the change having worked and consistent with it having done nothing, and only a controlled comparison separates those.
What This Protocol Cannot Establish
It measures what engines returned, on your prompts, on your dates, under your recorded conditions. That is all it measures.
It cannot establish why an engine did what it did. No provider documents how retrieved evidence becomes an ordered list of named vendors, so an outside observer is inferring mechanism from output, and an inference from output is not a mechanism.
It cannot attribute a change to your work without a controlled comparison. Observational rounds describe; they do not isolate.
It cannot tell you what buyers did. An AI-generated shortlist is a proposal a person may adopt, edit, ignore or check - and survey evidence indicates many do check. Presence in an answer is not presence in a consideration set.
And it cannot be attributed to a named model release. An outside observer often cannot tie a consumer response to a stable, externally verifiable model and retrieval version, which is why the observation date is the scope anchor on every figure a panel produces.
None of that makes the measurement pointless. It makes it a description of a moving quantity, which is what it is, and describing it honestly is more useful than a number with an unearned causal story attached.
Want a panel that will survive scrutiny?
Most of the work is in the decisions made before the first run - prompt selection, account state, coding rules - and those are quicker to get right with a second pair of eyes than to fix afterwards. Book a Session.
Sources
Sources: Platform documentation, quoted as read on August 31, 2026: Google Search Console Help, Generative AI performance report (Search) - on partial rollout, so its availability wording may change; Google Search Help, AI Mode in Google Search and Personal Intelligence in AI Mode; OpenAI Help Center, Searching the web with ChatGPT - the OpenAI article displays a relative update date rather than a fixed one, so it is cited to the review date. Measurement research: Schulte, Bleeker and Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv preprint, submitted April 8, 2026; not peer-reviewed; four engines, four commercial verticals, eight prompts per campaign, Swiss German-language market; the per-day run counts come from a simultaneous dataset and the two-to-four-week window from a separate temporal dataset); AirPulse (published July 20, 2026, updated August 3, 2026; frozen cohort of 30,504 response observations from 456 production jobs across four engines over a 30-day window - vendor research on an observational production cohort its authors state is not a random sample; it adopts the seven- and eight-run counts from the St. Gallen paper and states that its own cohort does not prove them universally sufficient, so the two sources are not independent derivations of the same thresholds); Żatuchin, "Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers" (arXiv preprint, July 14, 2026; not peer-reviewed; 12,933 responses, 20 Central and Eastern European brands, eight languages, three models - outcome measured is per-response sentiment polarity, not shortlist inclusion, with 91.9% of responses scoring exactly neutral and a recommendation-indicator refit named as future work); and Sharma, "The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries" (arXiv preprint, January 1, 2026; based on M.Tech thesis research at the Indian Institute of Technology Patna, per the arXiv listing; 112 Product Hunt-featured products, 2,240 queries, two model endpoints; the paper's own limitations note the Product Hunt sample skew and the two-model coverage). None of the four research sources above has been peer-reviewed, and none of the studies cited on this page measures B2B software shortlist stability directly. Their figures use different outcomes, samples, engines and dates and are not pooled anywhere on this page.
About the author
Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.