You can find out which domains AI systems cite. Several vendors publish rankings, some of them built on very large samples, and the broad shape is consistent enough to be useful: encyclopaedic references, community forums, video, professional networks and a long tail of publishers, with review platforms a meaningful minority in commercial queries. What you cannot do is put two of those published figures side by side. They count different things, out of different totals, on different engines, in different months, over prompt sets whose composition is not published. Two numbers that both look like "Reddit, about 11%" turn out to be a share of citations in one study and a share of responses in the other, five months apart.

And none of it measures influence. Every located study records what appeared in an answer. No located study measures what a cited source changed about the recommendation. This page sets out what has actually been measured, what each figure is a percentage of, what the providers document, and what remains open. It is the source-composition part of the vendor-shortlist pillar, which the pillar only summarises.

Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked platform documentation, quoted. Recommendation means a workflow that CoreAEX prescribes and no platform documents. Research findings are attributed in the prose with their sample, dates, engine and the outcome they measured, and are not tagged - they describe what was observed, not product documentation. Untagged text is ordinary explanation.

Four Different Things Get Published as One Number

Before any published figure is usable, establish what it counts and what it divides by. Four distinct quantities circulate under the word "citations," and they do not convert into one another.

  • Citation prevalence - the proportion of responses that cite a given domain at least once. Denominator: responses. A domain cited in half of all answers has 50% prevalence even if it is one link among twenty in each.
  • Citation share - a domain's citations divided by all citations recorded. Denominator: citations. A domain can hold a large share while appearing in a minority of answers, if it is cited repeatedly when it appears.
  • Share within a subset - a share computed over one class of link rather than all links. A platform holding 40% of review-platform links may hold 3% of all links.
  • Brand mention - the brand named in the text, with no link and no source attribution at all. Not a citation, and measured by different tools.

These are not pedantic distinctions. They produce numbers of very different magnitude from the same underlying answers, and the difference runs in no fixed direction, so there is no correction factor that turns one into another. One study on this page reports both a prevalence and a share for the same source class, from the same runs: review platforms appeared in 34.5% of responses and accounted for 8.5% of all links. A fourfold gap, same data, same day.

The practical rule follows directly. Record the metric and the denominator beside every figure you keep, and never compare two figures whose denominators differ. A composition number without its denominator is not a weak data point; it is not a data point. Recommendation

What the Composition Studies Actually Report

Two large, frequently referenced general-web studies illustrate why composition figures need their scope attached. They are not versions of the same measurement, and reading them side by side is the clearest demonstration of the problem.

Two frequently referenced general-web composition studies, with the outcome each reports
StudyMetric and denominatorEnginesWindowSample as stated
Semrush Prevalence - "% of responses" ChatGPT search, Google AI Mode, Perplexity Weekly snapshots, July 14 - October 12, 2025 "230K prompts"; "over 100M total AI citations"; each week "the top 25 domains"
Similarweb Share - "a percentage of total citation events in the measurement window" ChatGPT (web browsing mode), Google AI Mode January - February 2026 Per-domain counts and shares, United States, per engine

Semrush's design is thirteen weekly snapshots across three engines. Its published analysis characterises the top of the distribution: it states that it analysed over 100 million citations, and that each week, we studied the top 25 domains that appeared most often among AI citations. Domains outside a given week's top 25 are not represented in the published series, so the study cannot describe the long tail or its movement - which is a limit on what has been published, not a claim that the underlying data missed them. Its headline movements are large: ChatGPT cited Reddit in close to 60% of prompt responses in early August before collapsing to around 10% by mid-September, and Wikipedia dropped from appearing in roughly 55% of AI prompt responses to less than 20%. The authors state plainly that the data doesn't tell us exactly why these shifts happened - but it does show how quickly AI citation patterns can change. No geography is stated.

Similarweb's design is a share table per engine. It defines its unit precisely: a "citation event" is recorded each time ChatGPT (web browsing mode) references a domain in a response, drawn from citation events across monitored prompts in the United States, with citation share… a percentage of total citation events in the measurement window. On ChatGPT it reports wikipedia.org at 13.15% and reddit.com at 11.97% of citation events; on Google AI Mode, fandom.com at 7.16% and reddit.com at 4.19%.

Now put the two Reddit figures together, which is what most secondary coverage does. Semrush has ChatGPT citing Reddit in around 10% of responses by mid-September 2025. Similarweb has reddit.com at 11.97% of ChatGPT citation events in January and February 2026. The numbers look like agreement across five months. They are not comparable at all: one is the fraction of answers containing at least one Reddit link, the other is Reddit's slice of every link recorded. A domain can hold 12% of all citations while appearing in 40% of answers, or in 4%. Nothing in either study lets you infer the other's number.

There is a further limit both share. Each describes the prompt set its publisher monitors. Semrush states a prompt count and Similarweb states a geography and a tool, but neither publishes how its prompts were selected, what mix of intents or verticals they represent, or whether the set is weighted toward anything. A composition figure computed over a private prompt list is a description of that list. That is not a criticism of either study - it is the boundary on what the number can be used for, and it is the reason this page reports no single "most cited domains" ranking as though it were a property of the web.

Alongside those displayed-citation studies, a separate academic preprint is valuable for a different question: how far AI Overview source domains extend beyond Google's organic results. Kirsten and colleagues, at Ruhr University Bochum, UAR RC Trust and MPI-SWS, ran 4,706 queries across seven datasets in the US and Germany and found that 53% of the domains Google's AI Overviews consulted were absent from the organic top 10 for the same query, and 27% were absent from the top 100. That measures retrieval footprint - where consulted domains sit relative to organic results - not which domain classes dominate displayed citations, so it does not sit on the same axis as the two tables above. Its query mix is general-interest, not B2B software.

What the Providers Document About Source Preference

No provider documentation located names a preferred class of source. There is no documented tier for review platforms, analyst sites, vendor domains or community forums - not a statement that these are treated equally, but an absence of any published ranking among them.

Google's guidance for AI features states that eligibility rests on ordinary indexing and snippet eligibility, and that there are no additional technical requirements. Documented Its generative-AI optimization guide defines grounding, in its retrieval-augmented-generation section, as a technique that improves AI responses by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index, which pushes the question back into classic ranking rather than answering it with a source hierarchy. OpenAI's help article on searching the web describes query rewriting and follow-up queries, and states that ChatGPT ranks search results using multiple factors intended to help users find relevant, reliable information. Placement is not guaranteed., again without a source taxonomy.

Perplexity is the interesting case, because it publishes source labels and therefore has a taxonomy that can be inspected directly. It names three: Government, the official website of a government organization; Academic, scientific sites; and Trusted, a broad category for sites that appear often enough in results to receive a label and that publish information within their own area of expertise. The stated purpose is orientation for the reader - labels help you see at a glance what kind of source you are reading. There is no commercial category, no vendor category and no review-platform category. Documented A team hoping to find that review platforms or vendor documentation occupy a named tier will not find one there.

So the honest position on source preference is an absence with a defined scope: we read Google's AI-features and generative-AI documentation, OpenAI's search help article and Perplexity's source-label documentation in full, and none of them states a preference among source classes. That is different from saying source class does not matter. It means the question has not been answered by the people who would know, and what remains is third-party measurement of outputs.

What Has Been Measured for B2B Software Specifically

Three measurements come close to the B2B software question, and each answers a narrower question than its headline suggests.

Review-platform domains among cited articles. Citera analysed roughly 350,000 B2B SaaS articles across 10,382 keywords in 52 categories, querying ChatGPT with web search, Claude with web search, Perplexity and Google AI Overviews in May 2026 on the US market in English. It reports that 19% of AI-cited articles are hosted on B2B review platform domains (G2, Capterra, TrustRadius), compared to 9% for non-cited articles. It publishes a numbered methodology section and states its own limits without prompting, including We do not measure backlinks, Domain Rating, or any authority metric. All comparisons in this study are confounded by domain authority. and Each of the 10,382 keywords was queried once per AI engine. Both matter. The first means the 19-versus-9 gap cannot be separated from the fact that review platforms are large, established domains. The second means each data point is a single observation of a system that, as the variation evidence shows, moves substantially between runs.

Review platforms in commercial AI Overviews. SE Ranking analysed 30,000 keywords producing 22,729 generated AI Overviews, on Google AI Overviews only, in the United States, in a snapshot dated December 1, 2025. It reports that out of all AI Overview responses, 34.5% cited at least one review platform, while review platforms represent just 8.5% of all links. Its per-platform breakdown - the numbers most often quoted - consists of shares of review-platform links only, not of all links, so a platform reported at a large percentage there holds a small fraction of the whole. One engine, one country, one snapshot.

Owned versus external sources. Aleyda Solis analysed 15 leading brands across SaaS, ecommerce and finance on Google AI Mode, Gemini and ChatGPT, categorising each brand's top ten cited source domains as owned, social and community, news and review, competitor or other third party, and weighting them by mentions. In that mix, SaaS showed 17.7% owned against 82.3% external - the highest external share of the three verticals, ahead of finance at 79.4% and ecommerce at 69.6%. Two limits belong beside that number. The vertical figure rests on the SaaS portion of a 15-brand sample, not on 15 SaaS brands, and no geography is stated. Treat it as evidence that external sources can dominate a brand-level top-domain panel, not as an estimated rate for SaaS citations generally. Even scoped that tightly it is worth stating, because the direction is the one most often assumed backwards.

Read together, these three say something modest and useful. External sources occupy a substantial share of the displayed citation mix - and, in Solis's small top-domain benchmark, a majority of it. Review-platform domains are over-represented among Citera's AI-cited article group relative to its non-cited group, on a comparison the authors state is confounded. External is not the same as independent. The classes these studies code include social platforms, communities, review sites, competitors, marketplaces and other domains; some may carry independent evidence, but none of the studies codes independence, so the word does not belong on these figures. And none of the three establishes that any source class caused a recommendation. They record what was cited.

Why the Composition Keeps Moving

Composition is not a stable property you can measure once. Semrush's own thirteen-week series is the plainest demonstration in the published record: Reddit's prevalence in ChatGPT responses moving from close to 60% to around 10% inside about six weeks, with the authors explicitly declining to explain why. A ranking of most-cited domains is a photograph, and the subject moves.

That has two consequences for how a composition figure should be handled. The first is that every figure needs its window attached, not just its date of publication, because a study published in April may describe January and be revised in August. The second is that a change in composition between two of your own observations is not evidence about your content - the variation page covers how much movement is ordinary before any change is interpretable, and the measurement page covers the panel design that makes the question answerable at all.

The practical form: if you are tracking which sources appear on your own category's prompts, treat it as a repeated measurement on a fixed prompt panel, record the engine and conditions for every run, and report composition per engine as a distribution over the window rather than as a single ranking. Recommendation A one-off scrape of which domains an engine cited on your category last Tuesday is a real observation and a poor basis for a decision.

Figures This Page Does Not Use, and Why

Several widely-circulated composition numbers are absent from this page deliberately, and the reasons are analytical rather than procedural. Listing them is more useful than quietly omitting them, because they are the figures a reader is most likely to encounter elsewhere.

Aggregated cross-study indexes. One widely-shared ranking of the fifty most-cited websites states that it "synthesizes verified citation data from six studies conducted between August 2024 and April 2026" by taking "average rank position weighted by each study's citation volume." The component studies measure different outcomes over different windows on different engines. Averaging rank positions across incompatible denominators produces a number that has no denominator of its own, so there is no question it can answer precisely. The index does publish a methodology section, which is why the objection is to what that method does, not to its absence.

A large B2B SaaS citation dataset whose metric is undefined. One frequently-quoted analysis states a substantial sample - 57.2 million citations across 50 B2B brands and seven verticals, "across five major platforms over 60 days" - but does not define a citation, name the five platforms, give start and end dates or state a geography on the public page, and routes the full report through a download form. A very large sample does not compensate for an undefined unit, and a figure whose engines are unnamed cannot be scoped.

A citation study whose prompt set selects its own answer. Another analysis counts over 10,000 citations across "hundreds of high-intent SaaS prompts," selected as "prompts that include 'tools', 'software', 'alternatives to'." Prompts chosen for asking about tools and alternatives will disproportionately return list-format and comparison domains. The result is largely a property of the prompt selection, and the article states no exact prompt count, research-period dates, geography or significance testing against which that could be checked.

Vendor figures with untraceable attribution. A vendor report on page types and AI citations states that "pages with rich schema are 13% more likely to earn AI citations," credited only to the vendor's own unspecified research, with no sample, window or method given. A separate blog post from the same vendor, covering the same topic, mixes named external studies with its own unsourced internal figures the same way - a pattern worth watching for on this vendor's content generally, not just on one page. An untraceable number is not a weak number; it is one a reader cannot check.

And one reconciliation problem worth recording. As read on August 31, 2026, the Similarweb post used above describes "nearly 600,000 citation events… across ChatGPT and Google AI Mode," while its own per-engine counts and percentages imply denominators of roughly 899,500 for ChatGPT and 591,200 for Google AI Mode - each table internally consistent across its rows, and each individually above or near the stated combined total. The cause is not documented, and this page does not speculate about one. The handling that follows is narrow: use the shares within each engine's internally consistent table, and do not use or reconstruct a combined denominator until Similarweb clarifies the page. Living pages are re-read on the day you cite them, not on the day you found them.

What Is Not Known

No located study measures the source composition of B2B software recommendations with a defined metric, a published prompt set and repeated runs together. Each of the three closest measurements has one or two of those and not the third: Citera has a published methodology and a defined comparison but queried each keyword once; SE Ranking has a defined prevalence and share but one engine on one day; the two large general studies have scale and repetition but general-web prompt sets whose composition is not published. We searched for a B2B-software-specific composition study combining all three and located none.

Influence is not measured anywhere in this record. Every figure on this page describes what appeared in an answer. None of the located work varies a source and observes what changes in the recommendation, so nothing here supports a claim that citing, appearing on, or being reviewed by any particular class of source causes a vendor to be recommended. That experiment would need an intervention and a control, and no published study of B2B software recommendations has run one.

What each provider actually retrieves from, as opposed to what it displays, is not observable from outside. A displayed citation is evidence that a source was surfaced, not that it was decisive, and not that it was the only source consulted. Google's own documentation describes AI features as grounded in core ranking systems without publishing what the model read, and no provider exposes the retrieved set.

Of the sources used on this page, one is academic and the rest are vendor or practitioner research. None of the vendor studies is peer-reviewed, and each of those publishers sells software or consulting in search and AI visibility, which is why each is named as a vendor in the text above rather than presented as a neutral observer. The two general-web studies neither confirm nor contradict each other, because they have no outcome in common. Their figures are not pooled anywhere on this page and should not be pooled anywhere else.

Trying to work out which sources matter in your category?

The answer usually starts with running your own prompt panel rather than reading someone else's ranking - and that is a short conversation. Book a Session.

Sources

Sources: Platform documentation, quoted as read on August 31, 2026: Google Search Central, AI features and your website and Google's Guide to Optimizing for Generative AI Features on Google Search; OpenAI Help Center, Searching the web with ChatGPT - the OpenAI article displays a relative update date rather than a fixed one, so it is cited to the review date; Perplexity Help Center, Understanding source labels (last updated August 7, 2026). Academic research: Kirsten, Grosse Perdekamp, Wu, Upadhyay, Gummadi and Zafar, "Characterizing Web Search in the Age of Generative AI" (arXiv preprint v2, revised May 31, 2026, first submitted October 13, 2025; Ruhr University Bochum, UAR RC Trust and MPI-SWS; 4,706 queries across seven datasets, US and Germany - figures here are taken from v2 and the preprint has not been peer-reviewed). Vendor and practitioner research, each named as such: Semrush, "The Most-Cited Domains in AI: A 3-Month Study" (230,000 prompts across ChatGPT search, Google AI Mode and Perplexity; weekly snapshots July 14 - October 12, 2025, tracking each week's top 25 domains; outcome is percentage of responses; no geography stated); Similarweb, "Most Cited Domains in LLMs" (published April 14, 2026, last modified August 24, 2026; ChatGPT web browsing mode and Google AI Mode, monitored prompts in the United States, January - February 2026; outcome is citation share of total citation events, reported per engine - prompt selection and set composition are not published, and the stated overall total does not reconcile with the per-engine tables at the review date, as described above); Citera, "An Analysis of 350,000 B2B SaaS Articles" (10,382 keywords in 52 categories; ChatGPT and Claude with web search, Perplexity and Google AI Overviews; May 2026, US market, English; the authors state that every comparison is confounded by domain authority and that each keyword was queried once per engine); SE Ranking, "Despite 90% Traffic Loss, Review Platforms Top AI Overview Citations" (published January 29, 2026, last updated April 20, 2026; 30,000 keywords producing 22,729 generated AI Overviews, Google AI Overviews only, United States, data collected December 1, 2025; review platforms in 34.5% of responses and 8.5% of all links, with per-platform figures being shares of review-platform links only); and Aleyda Solis, "AI Search Is a 3rd-Party Citation Problem With an On-Page Corroboration Base" (published August 2, 2026, modified August 8, 2026; 15 leading brands across SaaS, ecommerce and finance on Google AI Mode, Gemini and ChatGPT; the metric is a mention-weighted mix of each brand's top ten cited source domains; the 17.7% owned against 82.3% external figure describes the SaaS portion of that 15-brand sample rather than 15 SaaS brands, and no geography is stated - reported here as a direction rather than a rate). Figures discussed in "Figures this page does not use" are described from their published pages, read in full on August 31, 2026, and are not relied on for any claim. No measurement cited here establishes that any source class causes a vendor recommendation, and figures from different studies are not pooled or compared numerically anywhere on this page.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.