AI-assisted research can influence which software vendors a buyer considers, and the systems doing it publish very little about how they choose. Google documents that AI Overviews and AI Mode may use a "query fan-out" technique and that their answers are grounded in its Search index. OpenAI documents that ChatGPT ranks search results using multiple factors intended to help users find relevant, reliable information. Placement is not guaranteed. Beyond statements at roughly that level of detail, the selection process is not public for any major engine. What a SaaS company can change is its evidence layer: whether its product facts are accessible, specific, current, consistent across the sources a system might read, and supported somewhere other than its own website. That improves the conditions under which a vendor can be considered. It does not control the result, and nothing located to date shows a page format, markup pattern, review profile or optimisation tactic that guarantees inclusion in an AI-generated shortlist.

"Earn a place" is used on this page in that sense and no other. It means improving the conditions for consideration. It does not mean a lever that produces inclusion.

What Counts as an AI-Generated Shortlist

Ten different things get reported as "AI visibility," and they have different evidence behind them, different measurement methods and different commercial value. Collapsing them is the most common analytical error in this subject, and it is expensive in a specific way: a team can celebrate a metric that has no relationship to whether a buyer ever evaluated the product.

Ten outcomes that are routinely reported as one
OutcomeWhat it meansWhat it does not establish
RetrievedA system fetched the page or pulled it into the context it was answering fromThat anything on the page reached the answer
MentionedThe vendor's name appears in the answer textThat the vendor was linked, endorsed or proposed for evaluation
CitedThe answer links or attributes information to a source URLThat the source supports what the answer says, or that the cited vendor is being recommended
RecommendedThe answer presents the vendor as a suitable option for the stated needThat the vendor appeared in a defined set the buyer treated as a shortlist
Shortlist inclusionThe vendor appears in the candidate set the answer proposesThat the buyer kept it
PositionWhere the vendor sits in the returned listA preference weight; ordering is unstable across runs (see below)
PersistenceThe vendor is still present after a follow-up questionStability across sessions, prompts, dates or users
ClickThe buyer opened the linkConsideration
Consideration setThe vendor is on the buyer's own evaluation listSelection
Selected or purchasedThe buyer boughtThat the AI answer caused it

Two of these deserve saying out loud, because they are the ones most often traded for each other in reporting. A cited vendor is not necessarily a recommended one - a citation can support a definition, a price, a criticism or a competitor's comparison table. And an AI-generated shortlist is not a buyer's shortlist. It is a proposal that a person may adopt, edit, ignore or check.

There is survey evidence that buyers do check. In the TrustRadius 2026 B2B Buying Disconnect Report (published July 15, 2026; 1,862 technology buyers and 444 vendors, fielded January 2026), 63% of buyers said they used AI during their purchase journey, and of those, 94% of buyers who used AI said they fact-check its responses at least some of the time. In a Gartner survey of 645 B2B buyers fielded August through September 2025 and released May 20, 2026, sixty-nine percent of B2B buyers prefer to validate AI-generated insights with sales reps, and buyers reported using an average of seven information sources during a recent purchase. Note the denominators: the TrustRadius figure is a share of AI users, not of all buyers, and the Gartner fielding predates the release by roughly eight months. Both are self-reported preference and behaviour, not observed conduct.

The practical reading is that an AI answer is one input among several into a decision that stays human, and the space it competes for is small. TrustRadius reports that 83% of buyers shortlisted three or fewer products, with an average shortlist of 2.7.

From One Buying Question to Several Research Tasks

Two providers document multi-query retrieval behaviour, in different words and for different products. Neither documents how retrieved evidence becomes a candidate vendor set or an ordered recommendation.

Google's AI features documentation (last updated December 10, 2025) states: Both AI Overviews and AI Mode may use a "query fan-out" technique - issuing multiple related searches across subtopics and data sources - to develop a response. Its generative-AI guide (last updated July 10, 2026) defines query fan-out as a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results to address the user's query, and describes the grounding step as relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index. Documented

OpenAI documents something analogous without using Google's label. Its ChatGPT search help article, under "Information shared with search providers," states that ChatGPT search typically rewrites your query into one or more targeted queries that it sends those providers, and that after reviewing the initial results, ChatGPT search may send additional, more specific queries to other search providers. Documented

Read those together carefully. They establish multi-query retrieval on two products, in two vocabularies, on two architectures that should not be assumed to be the same system. Neither describes how many sub-queries run, how their results are weighted, or how a set of retrieved pages becomes a list of named vendors in a particular order. No equivalent public mechanism description was located for Claude or Perplexity. Tools that claim to show you "your fan-out queries" are producing a model of the behaviour, not a readout of it.

The rest of this section is a conceptual model, offered because it is useful for planning coverage - not because any provider has published it. Take a realistic prompt: "best compliance platform for a 500-person fintech company." A buyer asking that has bundled several separable research needs into one sentence: what "compliance platform" means in fintech specifically, which products serve companies of roughly that headcount, which hold the certifications a regulated firm needs, what integrates with the systems that company already runs, what the implementation looks like, what it costs, and which are credible enough to defend internally. A system that decomposes the question is likely to need evidence on each of those, and that evidence lives in different places on a vendor's site - or nowhere.

The planning value of the model does not depend on it being architecturally accurate. If your product page answers the category question but nothing on your site states which company sizes you serve, which regulatory frameworks you support, or what integrates with what, then a research process that needs those facts cannot get them from you - regardless of how the process is implemented internally. That is a content-coverage argument, and it stands on its own. Recommendation

The Source Layers an AI Research System Can Draw On

A B2B software question can be answered from many classes of source, and we located no published research establishing a fixed hierarchy among them that holds across engines, categories and dates. What follows is a map of what exists, not a ranking of what wins.

Source classes available to an AI research system, and what a vendor controls in each
Source classTypically suppliesVendor control
Vendor website and product pagesCategory positioning, capabilities, use cases, claimed fitFull
Pricing pagesPrice, units, billing terms, tiers, or an explanation of why they are not publishedFull
Documentation and help centresFeature specifics, limits, configuration, integration detailFull
Security, compliance and integration pagesCertifications, data handling, hosting regions, connector listsFull, subject to what is true
Review platformsThird-party ratings, category placement, comparative languageIndirect - you can participate, not author
Customer evidenceNamed outcomes, deployment context, industry fitShared with the customer
Independent editorial comparisonsCategory framing, "best X for Y" lists, tradeoff judgementsNone directly
Communities and practitioner discussionUnfiltered experience, objections, migration storiesNone; participation is visible and reads as such
Analyst and research materialCategory definitions, market structure, vendor coverageCommercial relationship at most
Search indexes, feeds and provider-specific data sourcesWhatever the provider ingests and does not describe publiclyVaries; largely unobservable

Two boundaries govern how this table can be used. Source availability does not prove source use: a page being crawlable and indexed does not establish that any system read it while answering. And source use does not prove influence on a recommendation: a page can be retrieved and cited in support of a definition while playing no part in which vendors the answer proposes.

Providers do document some of this, unevenly. Google states that to be eligible as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements, and that there are no additional technical requirements. Documented OpenAI documents the input classes for its consumer shopping results - including structured metadata from first-party and third-party providers, and that product results are selected independently by ChatGPT and are not ads, nor influenced by any OpenAI partnerships - but that article addresses consumer merchandise and says nothing about software or B2B vendors, so it cannot be read across. Documented We reviewed Perplexity's public FAQ and located no statement describing how it selects, ranks or cites sources; that is a scoped absence in the material we read, not a demonstration that no such description exists anywhere.

One narrow question inside this table - whether a fact reaches an answer from a page's rendered text or from its markup - CoreAEX has begun testing directly. Round one was a URL-supplied content-access test, small and specific to pricing, run on August 30, 2026 against four purpose-built pages on a domain CoreAEX controls, logged out, from Serbia. On the markup-only page, the JSON-LD price appeared in 0 of 9 pricing responses across ChatGPT, Gemini and Google AI Mode. On the divergent page, eight of nine responses returned the visible $34, one returned no number, and none returned the JSON-LD value of $118. Claude and Perplexity produced no usable page data and failed every access check, so they are excluded from the denominators; content access for those two could not be established.

What that is, and what it is not. Separate access-check sessions for the same engine-page combinations returned distinctive visible-page details, which shows those engines could reach and read the pages during the observation window - it does not establish that each individual pricing response fetched the page. So this is an observation about which layer a price was reproduced from, on one date, on a domain with no standing in its category. It is not a measurement of vendor recommendation, shortlist inclusion or source preference; it does not establish why the engines behaved as they did; and it does not extend to any fact other than a price. Round two is scheduled for September 15, 2026 against the same unchanged pages, and will be reported as a second observation beside round one rather than pooled with it or substituted for it. The method, the denominators, the access limits and the full findings are in the AI SaaS pricing extraction study, part of the SaaS pricing schema cluster. The carry-over for this page is a single sentence: a commercial fact that exists only in markup should not be assumed to be extractable.

What Makes a Vendor a Defensible Candidate

Everything in this section is an evidence-quality and representation recommendation. None of it is a confirmed ranking factor, and it should not be sold internally as one. The argument for each item is that it makes a true, checkable statement about your product available where a research process - human or machine - would need it. That argument holds whether or not any given engine rewards it.

  • Explicit category and use-case positioning. A page that describes what the product does without saying what category it belongs to, or who it is for, leaves the matching work to inference. Name the category in the words buyers use, and name the situations the product is and is not built for.
  • Stated buyer constraints. Company size, industry, region, regulatory scope, team structure, deployment model. Constraint-heavy prompts are common in B2B; a vendor that does not state its constraints forces the system or the reader to infer them, or to rely on another source, and an inference can land either way.
  • Verifiable feature and integration information. Specific and checkable beats comprehensive and vague. A named integration list is a fact; "integrates with your stack" is not.
  • Current pricing, or an honest account of why it is not published. "Contact sales" is a legitimate model. An empty pricing page with no explanation of the model, the units or the range is a gap in your own evidence, and TrustRadius reports transparent pricing as buyers' top vendor wish-list item for four consecutive years.
  • Security and implementation evidence. Certifications with their scope and dates, hosting regions, data handling, and a realistic implementation description. In regulated categories this can become an early disqualifying filter.
  • Consistency across first-party and third-party sources. Where your site, your documentation, your review-platform profile and your partner listings disagree about a fact, a research process reading several of them at once encounters a contradiction, and you do not control which side it resolves to.
  • Comparisons that acknowledge tradeoffs. A comparison page that finds you superior on every axis is not usable evidence for anyone, including a reader.
  • Independent support for load-bearing claims. Claims that exist only on your own site rest on your own authority. The point is not that third-party sources are weighted more heavily - that is not established - but that a claim independently corroborated by a source with its own evidence is easier to verify. Repetition on its own can just as easily duplicate an error.

Recommendation for the block above.

One piece of research is worth naming here precisely because of what it does not license. Chu and Hou's preprint "Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems" (arXiv v2, revised August 21, 2026; first submitted June 16, 2026; not peer-reviewed) tested three commercial models on constructed sets of skincare products and reports that well-known brands were recommended 100% of the time when all products carried identical specifications, but that this dominance disappears with a rating advantage as small as 0.075 stars for a competitor (the paper's own abstract rounds this to less than a +0.1-star rating advantage). The same paper reports that authority-style marketing language, including fabricated clinical-evidence claims, broke that pattern. Two things follow and a third does not. It is consumer skincare, not B2B software, and it is a preprint on constructed product sets - so it is a hypothesis about a mechanism, not a finding about your category. Where it is genuinely useful is as a caution: a study in which fabricated evidence claims shifted model recommendations is a study describing an attack, not a tactic. Publishing claims you cannot substantiate is the thing this cluster exists to argue against, and the fact that it might work on a model in a controlled setting does not change that.

Why a Technically Visible Vendor Can Still Be Left Out

Absence from an answer is a symptom with many possible causes, and a single observation almost never identifies which one. The failure mode worth naming is the diagnostic shortcut: a vendor is missing from one answer, someone proposes a cause that matches a product they were already planning to buy, and the organisation acts on it.

Classes of explanation for an omission, and what would distinguish them
Possible causeWhat would have to be trueHow to narrow it
Poor match to the prompt's constraintsYour product genuinely does not fit the size, industry, region or requirement statedRun repeated matched panels with and without the constraint under the same recorded conditions. A persistent difference in inclusion rate is consistent with the constraint affecting selection; it does not show it was the only filter
Missing or ambiguous category positioningYour site never names the category the prompt usesSearch your own site for the prompt's category term
Insufficient supporting evidenceThe facts a comparison needs are absent or unspecific on your propertiesTry to answer the prompt yourself using only your own public pages
Inconsistent product factsYour site, docs and third-party profiles disagreeCompare the same five facts across all of them
Stronger evidence for competitorsCompetitors publish what you do notRun the same audit against two named competitors
Geographic or market limitationThe answer is region-scoped and you are not in that region's sourcesRepeat the same fixed panel in each location over the same dates, and report each location separately rather than as a difference
Retrieval or indexing differenceYour pages are not indexed, not fetchable, or not fetchable by that engine's agentCheck indexability, fetchability and bot-specific delivery as three separate things - the AI-crawler guide covers access and rendering. A request in a server, CDN or edge log shows that a request reached the logged layer; it does not show indexing, retrieval, content use, citation or recommendation
Answer variationYou appear in some runs and not othersRepeat the prompt; see the run counts in the next two sections
Provider-side selection behaviourSomething in the system's own process that is not publishedCannot be distinguished from outside. Record it as unexplained

The last row is not a formality. When the eight testable classes above have been checked and the omission survives, the honest conclusion is that the cause is unknown - not that the remaining explanation is whatever theory is currently fashionable. Recommendation

Why the Shortlist Changes Between Runs

Repeated asking of the same question does not return the same list, and the best available measurements put the instability high enough that a single observation carries very little information about a vendor's standing. This is the finding that most changes how a team should behave, and it is also the one most often overstated in the other direction.

One directly relevant volunteer-based practitioner experiment is by Rand Fishkin and Patrick O'Donnell, published January 28, 2026 and last modified June 28, 2026: 600 volunteers, 2,961 prompt runs across ChatGPT, Claude and Google's AI features, run in November and December 2025 on participants' own default settings. It measured exact-list and exact-order repeatability, and reports that ChatGPT and Google's AI features have a less than 1-in-100 chance of returning the same list of brands across two runs of the same prompt, and roughly 1 in 1,000 for the same list in the same order. The authors state plainly that they are not professional researchers or credentialed data scientists, that the work is not peer-reviewed, and that device, geography, login state and browsing history were not controlled.

A larger but differently designed study measures different outcomes, and the two should not be pooled. AirPulse (published July 20, 2026, updated August 3, 2026) analysed a frozen cohort of 30,504 valid response observations from 456 production jobs across ChatGPT, Gemini, Google AI and Perplexity over a 30-day window, with a separate 30,146-response validation cohort. Mention state flipped between consecutive observations 9.7% of the time; of the brand-prompt-engine cells observed at least three times, 35.9% contained both a mentioned and an unmentioned result within the window; and consecutive source-domain sets had a mean Jaccard overlap of 0.396. This is vendor research on an observational production cohort, and AirPulse states that it is not a random sample of all brands, prompts or AI users and does not establish causation. Exact shortlist identity, mention-state change and cited-source overlap are three different measurements. What the two studies jointly support is repeated, engine-separated measurement rather than a single screenshot.

The same study contains the correction to the nihilistic reading of it, and it matters more than the headline. While exact lists rarely repeated, individual well-established brands appeared very consistently: one healthcare provider appeared in 69 of 71 responses, one agency in 85 of 95, and leading headphone brands in 55-77% of responses. List instability and presence instability are different measurements. A vendor with strong, widely-distributed evidence can be present in most runs while never appearing in the same list twice.

Two preprints put numbers on what to do about it, and both numbers carry a scope. Schulte, Bleeker and Kaufmann of the University of St. Gallen, in "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv preprint submitted April 8, 2026; not peer-reviewed), argue that visibility should be characterised as a distribution rather than a single-point outcome. They recommend at least 7 runs per prompt per day for brand visibility monitoring and 8 runs when source-level coverage matters. Those counts come from a simultaneous dataset of up to ten repetitions per engine-prompt group. From a separate temporal dataset that tracked daily results over 45 to 46 days, they add that rolling aggregation over two to four weeks is therefore recommended to obtain per-brand estimates that are both statistically stable and representative of sustained visibility rather than momentary snapshots. The two numbers answer two different questions: how many runs to take on a day, and over how many days to aggregate them. The study covers four AI search engines, four commercial verticals, eight prompts per campaign and a Swiss, German-language market, and the authors note that results may not generalise to other regional or linguistic markets. Treat seven and eight as a starting floor from one study, not a universal sufficiency threshold - and raise them when the prompt is commercially important, when results sit near a decision threshold, when engines disagree, or when the set of sources is still moving. Separately, Dmitrij Żatuchin's variance-components decomposition (arXiv preprint submitted July 14, 2026; not peer-reviewed) analysed 12,933 responses across 20 Central and Eastern European brands, eight languages and three models, and found query language to be the largest systematic source of variance at 26.5% against 1.5% for brand identity - concluding that a single response carries almost no brand-discriminating signal, and that adding languages and models buys far more reliability per unit of budget than adding repetitions past the fifth. Scope that finding carefully: its outcome variable was per-response sentiment polarity toward a brand, not shortlist inclusion - 91.9% of responses in that dataset scored as exactly neutral, and the paper names a recommendation- or rank-based outcome as a proposed future extension it did not measure - so it speaks to how noisy this class of measurement is rather than to shortlisting specifically.

Documented sources of variation sit alongside the statistical ones. Google states that AI Mode doesn't always get it right and may misinterpret web content or miss context, and that in some cases, AI Mode will provide a set of web links if there's not high enough confidence in the quality or helpfulness of an AI response. Documented Google also documents personalisation, and the boundary inside it is worth keeping straight. Personal Intelligence in AI Mode references previous searches and activity saved in your Search Services History to provide suggestions tailored to your tastes and preferences when the relevant settings are enabled, and can be opted into or out of at any time. Google separately describes an AI Mode memory experience for details a user asks it to remember or forget; it is that narrower experience which the page states is available in the US, in English at the review date. Documented Prompt wording, follow-up turns, engine and product, model generation, geography, account state and the sources available on the day all move the answer, and several of them are not observable from outside. The size of that movement at four different time scales, the personalisation mechanisms Google and OpenAI document in full, and how to read a change in your own numbers without mistaking noise for signal are set out in why AI vendor shortlists change across prompts, buyers and markets.

How to Establish Where You Actually Stand

The unit of measurement is a prompt panel run repeatedly under recorded conditions, not a screenshot. The framework below is CoreAEX's, assembled from the run-count guidance in the two preprints above and from ordinary experimental hygiene. No platform prescribes it. It covers shortlist standing specifically; general citation measurement - what counts as a citation, which data sources can answer which question, and how to report a before-and-after honestly - is set out in measuring AI citations and is not re-derived here. Recommendation

  1. Fix a prompt panel and stop editing it. A panel that changes between rounds cannot show change over time. Write the prompts once, version them, and treat any edit as the start of a new baseline.
  2. Separate the prompt types, because they measure different things: category prompts ("best X platform"), use-case prompts ("X for onboarding contractors"), comparison prompts ("A vs B"), and constraint-heavy prompts ("for a 500-person fintech in the EU"). A vendor can be strong in one type and absent from another, and averaging them hides it.
  3. Run each prompt several times, on several dates. The published guidance available is seven or more runs per prompt per day, aggregated over a rolling two-to-four-week window - two separate numbers, from two datasets in one study in one market. Treat both as a starting floor rather than a settled standard, raise them where the decision matters, and record what you actually did.
  4. Record the conditions on every run: engine and product surface, model version if the interface exposes one, geography, language, account state, whether browsing or grounding was active, and the UTC timestamp. Several consumer interfaces expose no model version to logged-out users; record that as "not exposed" rather than leaving it blank or guessing.
  5. Code the outcomes separately, using the ten-way distinction from the first section: mentioned, recommended, included in the proposed set, position, cited, persisted after a follow-up. One column each. Do not collapse them into a score before the raw coding exists.
  6. Inspect the sources the answer actually displayed, and open them. This is the step teams skip and it is the one that pays: it tells you which of your properties, and whose third-party pages, are in play.
  7. Compare the answer against what those sources say. Where an answer states a fact about your product, check it against the page it cited and against your current truth. Wrong facts about you are a different problem from absence, and they need a different fix.
  8. Treat every change as an observation until you have controlled something. Between two rounds, the engine changed, your competitors changed, the web changed and you changed. A movement in your numbers after a content change is consistent with the change having worked and consistent with it having done nothing.

What to Fix First

The sequence below is ordered by how much of it is under your control and how confidently the work can be justified without appealing to a mechanism nobody has established. Every item is defensible as ordinary content and product-marketing quality, which is the point: none of it is wasted if the engines change. Recommendation

  1. Fix product information that is inaccessible, contradictory or out of date. This needs no theory of AI behaviour to justify. A fact that is wrong on your own site is wrong for every reader and every system.
  2. Make category and use-case fit explicit in the language buyers use, including the constraints you do and do not serve.
  3. Strengthen decision-stage evidence - integrations, security and compliance detail, implementation reality, pricing or a clear account of the pricing model.
  4. Align the load-bearing facts across your own properties - site, docs, help centre, partner and marketplace listings. Pick the five facts a buyer would disqualify you on and make them identical everywhere.
  5. Identify material third-party gaps. Where a category's evidence lives largely off your site - review platforms, editorial comparisons, community discussion - a missing or stale profile is a gap in the record, whatever any engine does with it.
  6. Establish repeatable measurement before making claims about improvement, using the panel above.
  7. Test changes rather than declaring success from one favourable answer. A single good output after a content change is very weak evidence, and it is the evidence most often presented.

What Is Not Known

Six boundaries, stated plainly, because each is routinely crossed in writing on this subject.

The selection processes are not public. Google documents the fan-out technique and its grounding in Search; OpenAI documents query rewriting, follow-up searches, and that ChatGPT ranks results using "multiple factors" with "placement is not guaranteed"; and we located no comparable published description of source selection or vendor-list construction from Perplexity's public FAQ. Those disclosures describe retrieval behaviour. What weighting any engine applies, and how a set of retrieved pages becomes an ordered list of named vendors, is documented by none of them.

Observed outputs do not reveal the algorithm. A pattern across answers is a pattern across answers. Inferring a ranking system from it is the same error as inferring Google's ranking algorithm from a page of results, with less stability in the data and no public documentation to check the inference against.

Buyer-adoption surveys do not establish revenue causality. G2's "The Answer Economy" (published April 15, 2026; 1,076 B2B decision-makers across North America, EMEA and APAC, fielded March 2026) reports that 51% of B2B software buyers now start their research with an AI chatbot more often than with Google and that 69% of buyers reported that chatbot guidance led them to select a different vendor than they had initially planned. Both are self-reported, and G2 is an interested party in the finding. The 51% figure is a comparison of frequency against Google, not a statement that half of buyers have replaced Google - the same report finds that 80% still use Google somewhere in their buying journey, treating it as complementary rather than essential. And G2 is a software review platform publishing research about the channel it competes in and sells into.

Survey findings on this topic are less stable than they look. G2's own 2026 Buyer Behavior Report (1,038 B2B decision-makers, fielded June 2026) reports that the top sources influencing buyer shortlist decisions are review sites (38%) and AI chatbots (37%) - a reversal of the ordering in its own April report, where generative AI chatbots led at 54% against 43% for software review sites. Different samples, different instruments, fielded roughly three months apart, same publisher. Treat neither ordering as established, and do not build a strategy on the gap between them. The one thing both support is that AI chatbots are now among the sources buyers say shape shortlists.

An AI recommendation is not a purchase. The closest behavioural measurement located is a vendor-authored, non-peer-reviewed preprint: Iannelli and Ai of Scrunch AI, "From Prompt to Purchase" (arXiv, submitted June 9, 2026). It links ChatGPT, Claude and Gemini conversations from an opt-in clickstream panel in two English-speaking markets to later web activity, for a curated lexicon of high-recognition consumer brands. Among users with no recent observed engagement, a recommendation was associated with a 4.3 percentage-point increase in same-name Google search and a 2.4-point increase in visits to the brand's own site, against matched backward placebos.

Four boundaries travel with that result. Google's search-embedded surfaces are excluded by design, because there a same-name search outcome would be mechanical. The design is observational. The authors state that they do not observe transactions - the retail measure is a purchase-adjacent retailer page visit. And they do not disclose aggregate panel size, per-analysis user counts or per-cell sample sizes, citing commercial sensitivity - though the paper adds that every headline cell clears its own minimum-disclosure threshold "by a wide margin." The authors are both affiliated with Scrunch AI, a vendor selling AI-visibility measurement, and the paper carries no separate conflict-of-interest disclosure beyond stating the measurement system runs in production there. The authors read the effect as upper-funnel brand exposure rather than acceleration of an existing purchase; nothing in it establishes the same effect for B2B software.

Being cited is not being represented accurately. The Tow Center for Digital Journalism's November 27, 2024 study asked ChatGPT to identify the source of 200 block quotes from 20 publishers and found partially or entirely incorrect responses 153 times, with the tool rarely signalling uncertainty; OpenAI disputed the test's representativeness. That is an old date in this field and a publisher-attribution task rather than a vendor-recommendation one, so it is not a current measurement of anything on this page - it is the clearest published demonstration that a citation and a correct statement are separate things.

What This Page Claims, and What It Does Not

AI-assisted research can influence which vendors enter a B2B software buyer's consideration set, and buyers report using it and then checking it. SaaS teams cannot control an AI-generated shortlist. They can make their inclusion more defensible by publishing accessible, specific, current, internally consistent and independently supported product evidence - work that is justified by ordinary content quality regardless of what any engine does with it.

What this page does not claim: that any tactic guarantees inclusion; that the engines behave alike; that a mention is a citation, a citation a recommendation, or a recommendation a decision; that a single answer measures anything; or that the movement in a metric after a content change demonstrates that the change caused it.

Want to know where your product actually stands?

The first useful step is usually not a tool - it is trying to answer your buyers' hardest constraint question using only your own public pages, and seeing where you run out of evidence. That is a short conversation. Book a Session.

Sources

Sources: Platform documentation, quoted as read on August 30, 2026: Google Search Central, AI features and your website (last updated December 10, 2025) and Optimizing for generative AI (last updated July 10, 2026); Google Search Help, AI Mode in Google Search and Personal Intelligence in AI Mode; OpenAI Help Center, Searching the web with ChatGPT and Shopping with ChatGPT Search - the two OpenAI articles display relative update dates rather than fixed ones, so they are cited to the review date above. Perplexity's public FAQ was reviewed on the same date and contained no description of source selection or ranking. Buyer research, all self-reported survey data: G2, "The Answer Economy" and its accompanying announcement (published April 15, 2026; n = 1,076 B2B decision-makers; fielded March 2026; North America, EMEA and APAC) and 2026 Buyer Behavior Report (n = 1,038; fielded June 2026) - G2 is a software review platform reporting on a channel in which it competes; TrustRadius 2026 B2B Buying Disconnect Report (published July 15, 2026; n = 1,862 buyers and 444 vendors; fielded January 2026) - also a review platform; and Gartner (released May 20, 2026; n = 645 B2B buyers; fielded August-September 2025). Their figures use different questions, denominators and field dates and are not pooled anywhere on this page. Measurement and model-behaviour research: Fishkin and O'Donnell, "AIs are highly inconsistent when recommending brands or products" (SparkToro with Gumshoe.ai, published January 28, 2026, modified June 28, 2026; 600 volunteers, 2,961 runs, November-December 2025 - practitioner research, self-described as not peer-reviewed); AirPulse (published July 20, 2026, updated August 3, 2026; frozen cohort of 30,504 response observations from 456 production jobs over a 30-day window, plus a 30,146-response validation cohort - vendor research on an observational production cohort, stated by its authors not to be a random sample); Schulte, Bleeker and Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv preprint, submitted April 8, 2026; four engines, four verticals, eight prompts per campaign, Swiss German-language market; run counts from a simultaneous dataset, the aggregation window from a separate temporal dataset); Żatuchin, "Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers" (arXiv preprint, July 14, 2026; outcome measured is per-response sentiment polarity, not shortlist inclusion); Chu and Hou, "Incumbent Advantage" (arXiv preprint v2, revised August 21, 2026, first submitted June 16, 2026; constructed consumer-skincare tests, three commercial models); Iannelli and Ai, "From Prompt to Purchase" (arXiv preprint, June 9, 2026 - authored at Scrunch AI, a vendor in this category; opt-in clickstream panel, two English-speaking markets, high-recognition consumer brands, Google AI surfaces excluded, no transactions observed, panel and cell sizes undisclosed); and Jazżwińska and Chandrasekar for the Tow Center for Digital Journalism, "How ChatGPT Search (Mis)represents Publisher Content" (November 27, 2024; 200 quotes, 20 publishers; disputed by OpenAI). None of the four preprints has been peer-reviewed. The CoreAEX source-layer figures are from round one of an ongoing first-party test, a URL-supplied content-access test run August 30, 2026 on four purpose-built pages; the method, denominators, access limits and full findings are in Which Layer Does an AI Answer Read a Price From? A Controlled Test.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.