Ask the same AI system the same software question twice and you will usually get a different list. In one directly relevant volunteer-based practitioner study, identical brand lists repeated in fewer than one run in a hundred. Across four engines measured daily for a month and a half, roughly two thirds of the cited sources changed from one day to the next. Two months apart, the links Google's AI Overviews returned overlapped 18% with the earlier run, against 45% for Google's organic results. But the thing that moves most is the composition of the list, not whether a well-established vendor appears in it at all - in the same volunteer study, individual strong brands showed up in 69 of 71 responses and 85 of 95. Variation is large, it is measured, and it is not evenly distributed across the questions a SaaS team actually cares about.

This page sets out how much moves, over what intervals, what is documented about why, and how to tell an actual change in your position from ordinary noise. It is the part of the vendor-shortlist pillar that the pillar only summarises.

Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked platform documentation, quoted. Recommendation means a workflow or threshold that CoreAEX prescribes and no platform documents. Research findings are attributed in the prose with their sample, dates and the outcome they measured, and are not tagged - they describe model behaviour, not product documentation. Untagged text is ordinary explanation.

Two Different Things Get Called "My Shortlist Changed"

Before any number is useful, separate list identity from presence. They are different measurements with very different stability, and a team that merges them will conclude the whole channel is random when its own position may be stable.

  • List identity - the exact set of vendors returned, and their order. This is highly unstable.
  • Presence - whether your vendor appears anywhere in the answer at all. This is considerably more stable, particularly for established brands.

The clearest direct illustration comes from one uncontrolled volunteer dataset, with a second vendor dataset supporting the same direction on a different metric. Rand Fishkin and Patrick O'Donnell's volunteer study, published January 28, 2026 and last modified June 28, 2026, ran 2,961 prompts across ChatGPT, Claude and Google's AI features with 600 volunteers on their own default settings in November and December 2025. It reports a less than 1 in 100 chance of the same list of brands recurring, and roughly 1 in 1,000 for the same list in the same order. In the same runs, one healthcare provider appeared in 69 of 71 responses, one agency in 85 of 95, and leading headphone brands in 55-77%. The authors state that they are not professional researchers or credentialed data scientists, that the work is not peer-reviewed, and that device, geography, login state and browsing history were not controlled.

A second dataset points the same way on a different metric. AirPulse, an AI-visibility vendor, published a frozen cohort of 30,504 response observations from 456 production jobs across ChatGPT, Gemini, Google AI and Perplexity over a 30-day window (published July 20, 2026, updated August 3, 2026). Mention state flipped between consecutive observations 9.7% of the time. Of the brand-prompt-engine cells observed at least three times, 35.9% contained both a mentioned and an unmentioned result - which means the complement, roughly 64%, were consistently one or the other throughout the window. AirPulse states the cohort is not a random sample of all brands, prompts or AI users and does not establish causation.

So the practical reading is narrower than "AI answers are random." The set is unstable; membership is more stable than the set. Which of those two you are measuring decides whether your numbers mean anything.

How Much Changes, at Four Time Scales

Instability is not one number. It depends entirely on the gap between the two observations you are comparing, and the published measurements span four different intervals with four different outcomes. They are not interchangeable and they are not pooled here.

Measured instability by interval, with the outcome each study reported
IntervalMeasuredOutcome variableSource and scope
Same prompt, repeated within 24 hours Pairwise Jaccard similarity of cited sources averaged 0.32-0.43 across four campaigns Overlap of the cited-source set St. Gallen preprint; ChatGPT, Perplexity, Gemini, Google AI Mode; up to 10 repetitions per prompt; March 21-25, 2026; Swiss, German-language
One day to the next Day-to-day Jaccard similarity of cited sources averaged 0.34-0.42 Overlap of the cited-source set Same study; 45-46-day window, January 24 - March 20, 2026
Between consecutive observations of the same cell Mention state flipped 9.7% of the time; mean Jaccard overlap of consecutive source-domain sets 0.396 Whether the brand was mentioned; source-domain overlap AirPulse; 30,504 observations; four engines; 30-day window
Two months apart Jaccard similarity of retrieved links 18% for Google AI Overviews, against 45% for Google organic Overlap of the retrieved-link set Kirsten et al. (arXiv v2, revised May 31, 2026; now peer-reviewed, ACL 2026 Findings); 4,606 repeated queries across six datasets - the seventh, 100-query Trends dataset was collected only in September and was not part of the temporal comparison; US and Germany, English-language queries; July/August 2025 repeated in September 2025

Two things in that table deserve to be read carefully rather than skimmed.

First, the within-day figure is barely better than the day-to-day figure. St. Gallen's "Don't Measure Once" (arXiv preprint submitted April 8, 2026; not peer-reviewed) reports 0.32-0.43 for repetitions inside a 24-hour window and 0.34-0.42 across consecutive days. The authors gloss the day-to-day number plainly: A Jaccard value of 0.35 implies that, on average, only about 35% of the cited sources overlap between two consecutive days - meaning roughly 65% of all sources change from one day to the next. That the two intervals produce similar figures indicates that substantial churn exists even over short windows and cannot be explained only by day-to-day changes in the public web. It does not apportion the remaining causes: the design does not isolate provider-side, index or retrieval changes, and the within-24-hour collection itself ran across five dates. What it does establish is narrower and still useful - waiting a day does not make a single observation representative.

Second, the two-month comparison is the one with a control. Kirsten, Grosse Perdekamp, Wu, Upadhyay, Gummadi and Zafar (arXiv preprint v2, revised May 31, 2026, first submitted October 13, 2025; since accepted and published in Findings of the Association for Computational Linguistics: ACL 2026, so no longer accurately described as unreviewed) repeated their experiments roughly two months apart and measured both a generative surface and a conventional one on the same queries. Comparing the Jaccard similarity between sets of retrieved links across runs, they report the highest overlap for Organic search (45%) while AIO exhibits substantially lower overlap (18%). That is the useful shape of the finding - not that generative answers churn, but that they churn markedly more than the ranked results on the same engine, over the same interval, for the same queries. The comparison covers the six datasets collected in both periods - 4,606 queries - since the seventh, a 100-query Google Trends set, was collected only in September. Their scope is US and Germany, with all queries issued in English, on a query mix weighted toward consumer and general-interest topics rather than B2B software.

All four rows measure source or mention overlap. None of them measures shortlist inclusion for B2B software vendors specifically, and none should be quoted as though it did.

What Moves the Answer

Nine things plausibly move an AI-generated vendor list, and they differ sharply in how well established they are and in whether you can do anything about them. The table separates those two questions, because conflating them is how teams end up trying to control the uncontrollable and ignoring the parts they own.

Sources of variation, their evidence status and your control over them
Source of variationEvidence statusUnder your control
Resampling - the same prompt, same conditions, run againMeasured. The largest single component in the one variance decomposition locatedNo. Only your measurement design responds to it
Prompt phrasingMeasured indirectly; distinct prompts produce distinct source setsNo, but you choose which phrasings you track
Query languageMeasured. The largest systematic component in that same decompositionOnly by choosing which markets you publish for
Engine and product surfaceMeasured across every study here; the four engines disagree consistentlyNo
Model generation and provider-side changeInferred. Consumer interfaces expose no version string to logged-out users, so this cannot be isolated from outsideNo
GeographyDocumented by OpenAI as a possible input to ChatGPT search; measured as a study condition in cross-market researchNo, but it is recordable
Account state, Search Services History, memory and personalisationDocumented by Google and OpenAI (below)No
The web changing - new pages, edits, removalsPartly separable: within-day churn is nearly as high as day-to-day churn, so this explains less than it appears toOnly your own properties
Your own evidence changingThe one input you control, and the hardest to detect against the noise aboveYes

The variance decomposition referred to twice above is Dmitrij Żatuchin's crossed random-effects analysis (arXiv preprint submitted July 14, 2026; not peer-reviewed), covering 12,933 responses across 20 Central and Eastern European brands, eight languages and three models. It attributes 34.8% of the variance, on the stability subset it analyses, to pure within-prompt resampling, and 26.5% to query language, against 1.5% for brand identity - concluding that a single response carries almost no brand-discriminating signal. Its outcome variable is per-response sentiment polarity toward a brand, not shortlist inclusion, so it establishes how noisy this class of measurement is rather than how vendor lists specifically behave. Read as a statement about measurement, it is directly useful. Read as a statement about shortlisting, it is a substitution.

The Personalisation That Providers Document

Two providers document mechanisms by which two buyers asking the same question receive different answers. This is worth stating precisely, because it is one of the few places where the reason for variation is on the record rather than inferred. Everything quoted in the three paragraphs below is published platform documentation. Documented

Google documents that Personal Intelligence in AI Mode references previous searches and activity saved in your Search Services History to provide suggestions tailored to your tastes and preferences when the relevant settings are enabled, and can be opted into or out of at any time. It separately describes an AI Mode memory experience for details a user asks it to remember or forget; that narrower experience is the one the page states is available in the US, in English at the review date.

OpenAI documents an equivalent for ChatGPT search: If memory is enabled, ChatGPT may use relevant saved memories when rewriting a search query. The example given is a user who has shared that they are vegan and live in San Francisco receiving a search for "good vegan restaurants San Francisco." The same article states that a VPN or network location may affect the approximate location inferred from your IP address and that if memory is enabled, saved location information may also influence search results.

Google also documents the limits of the output itself. Its AI Mode help page states that AI Mode doesn't always get it right and may misinterpret web content or miss context, and that in some cases, AI Mode will provide a set of web links if there's not high enough confidence in the quality or helpfulness of an AI response. That last sentence describes a documented branch in the product: for some queries on some occasions, there is no generated answer to be included in.

The practical consequence is a constraint on measurement rather than a lever. A logged-in account with memory enabled is a different measurement instrument from a fresh logged-out session, and neither is more correct - they answer different questions. Pick one, record it on every run, and do not compare across it. Recommendation

Reading a Change in Your Own Numbers

Against churn of this size, a single before-and-after comparison cannot distinguish an improvement from an ordinary fluctuation. That is not a counsel of despair; it is a statement about how many observations a claim needs before it is worth making.

Three failure modes recur, and all three are avoidable.

The single-screenshot conclusion. One answer, on one day, in one session, is an observation with a very wide error bar around it. The St. Gallen authors argue that visibility should be characterised as a distribution rather than a single-point outcome. They recommend at least seven runs per prompt per day for brand-visibility monitoring and at least eight per prompt per day where source-level coverage matters - counts drawn from a simultaneous dataset of up to ten repetitions per engine-prompt group. From a separate temporal dataset tracked daily over 45 to 46 days, they add that rolling aggregation over two to four weeks is therefore recommended. Both come from one study covering four engines, four commercial verticals and a Swiss German-language market, and the authors note results may not generalise to other regional or linguistic markets.

AirPulse operationalises the same schedule for practitioners, explicitly using the St. Gallen paper as the source of the seven- and eight-run starting counts: Use seven brand-detection runs and eight source-coverage runs per prompt-engine cell as a starting floor, and Spread the runs across two to four weeks. Running eight times in ten minutes measures short-term generation variability. It does not capture day-to-day retrieval, model or index changes. Its own production cohort independently supports the broader need for repeated, time-distributed measurement, but it does not independently derive or validate those thresholds - AirPulse says so itself, stating that its study supports the direction - repeat measurements - but does not independently prove that seven and eight are universally sufficient. Two sources support the operating direction, but they are not independent derivations of the same thresholds. Treat the numbers as a starting floor.

The post-hoc explanation. A movement appears, and a cause is proposed after the fact that happens to match work the team recently did. With roughly two thirds of cited sources changing between consecutive days before anyone intervenes, a movement that follows a content change is consistent with the change having worked and equally consistent with it having done nothing. Attribution needs a design set up in advance - a fixed panel, repeated runs, recorded conditions, and a comparison that was specified before the result was seen. Recommendation

Comparing across a condition you changed without noticing. A different engine, a different account state, a colleague running the prompt from another country, a slightly reworded prompt, a logged-in session where the last one was logged out. Each of those is a documented or measured source of variation in the table above, and any of them can produce a difference that has nothing to do with your product.

The general principle behind all three: treat every change as an observation until something has been controlled. Between two rounds, the engines changed, your competitors changed, the web changed, and you changed. Only the last of those was yours.

What Variation Does Not Mean

Three readings of this evidence go too far, and the second is the one that costs money.

It does not mean AI answers are random. Random would mean no structure. What the data shows is structure with heavy noise on top: established brands appeared in 69 of 71 and 85 of 95 responses in the SparkToro runs, and roughly 64% of AirPulse's repeatedly-observed cells were consistently mentioned or consistently absent. Both of those are the opposite of randomness.

It does not mean measurement is pointless. A single response records what happened once, which is a real observation; what it cannot do is estimate a stable mention, citation or shortlist-inclusion rate. Repeated, condition-controlled measurement is what turns observations into a distribution. The distinction matters commercially, because the response to "this is unmeasurable" is usually to stop looking, which leaves a team with no view of its own position at all.

It does not mean your content work had no effect. A flat or noisy result after a content change is not evidence that the change did nothing - it is an absence of evidence either way, at that sample size. Publishing accurate, specific, current product evidence is justified by ordinary content quality and by readers who use it, whatever any engine does with it. That case does not depend on a measurable citation lift, and it survives an engine update.

What Is Not Known

No study located measures shortlist stability for B2B software vendors specifically. The measurements above cover cited-source overlap, source-domain overlap, mention state and retrieved-link overlap, on query mixes that are consumer-weighted, Central European, or unspecified in vertical. We searched for a B2B-software-specific stability measurement and located none with both a published methodology and a defined metric.

Why the churn happens is not established from outside. Resampling, index change, personalisation, provider-side model updates and prompt sensitivity are all plausible contributors, and the studies that can separate any of them do so partially. An outside observer often cannot attach a consumer response to a stable, externally verifiable model and retrieval version, so a change in behaviour cannot reliably be attributed to a named model release - which is why the observation date, not a model name, is the scope anchor on every figure here.

Of the five measurement sources used on this page, two are practitioner or vendor research, one - the Kirsten et al. study - is now peer-reviewed (published in ACL 2026 Findings after this page's sources were first drafted), and two remain unreviewed preprints (St. Gallen and Żatuchin). They agree that single observations are unstable, which is worth something; they use different outcomes, samples, engines and dates, which means their numbers are not pooled anywhere on this page and should not be pooled anywhere else.

Not sure whether your numbers moved or just wobbled?

The question is usually settled by how the measurement was set up, not by the numbers themselves - and that is a short conversation. Book a Session.

Sources

Sources: Platform documentation, quoted as read on August 30, 2026: Google Search Help, AI Mode in Google Search and Personal Intelligence in AI Mode; OpenAI Help Center, Searching the web with ChatGPT - the OpenAI article displays a relative update date rather than a fixed one, so it is cited to the review date. Measurement research: Fishkin and O'Donnell, "AIs are highly inconsistent when recommending brands or products" (SparkToro with Gumshoe.ai, published January 28, 2026, modified June 28, 2026; 600 volunteers, 2,961 runs, November-December 2025 - practitioner research, self-described as not peer-reviewed, with device, geography and login state uncontrolled); AirPulse (published July 20, 2026, updated August 3, 2026; frozen cohort of 30,504 response observations from 456 production jobs across four engines over a 30-day window, plus a 30,146-response validation cohort - vendor research on an observational production cohort that its authors state is not a random sample); Schulte, Bleeker and Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv preprint, submitted April 8, 2026; not peer-reviewed; four engines, four commercial verticals, eight prompts per campaign, Swiss German-language market; the per-day run counts come from a simultaneous dataset, the two-to-four-week aggregation window from a separate temporal dataset; AirPulse adopts the same window in its practitioner guidance while stating that its production cohort supports repeated measurement but does not independently prove the thresholds universally sufficient); Kirsten, Grosse Perdekamp, Wu, Upadhyay, Gummadi and Zafar, "Characterizing Web Search in the Age of Generative AI" (arXiv preprint v2, revised May 31, 2026, first submitted October 13, 2025; Ruhr University Bochum, the UA Ruhr Research Center for Trustworthy Data Science and Security, and the Max Planck Institute for Software Systems; since accepted and published in Findings of the Association for Computational Linguistics: ACL 2026 - no longer an unreviewed preprint as of this page's publication; 4,706 queries across seven datasets overall, of which the temporal comparison used the 4,606 queries from the six datasets collected in both periods; US and Germany, all queries in English; July/August 2025 with a September 2025 repeat - figures on this page are taken from the arXiv v2 text); and Żatuchin, "Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers" (arXiv preprint, July 14, 2026; not peer-reviewed; 12,933 responses, 20 Central and Eastern European brands, eight languages, three models - outcome measured is per-response sentiment polarity, not shortlist inclusion). Of the five sources above, the Kirsten et al. study is peer-reviewed; the Schulte/Bleeker/Kaufmann and Żatuchin preprints are not; none of the measurements cited on this page directly measures B2B software shortlist stability. Their figures use different outcomes, samples, engines and date ranges and are not pooled anywhere on this page.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.