Review platforms are a recurring but minority presence in what AI answers cite, and the honest answer to the title question is that nobody has tested it. No located study varies a review profile and measures what changes in a recommendation. What exists is observational: counts of how often review domains appear, and one published regression on whether review volume predicts citations. That regression found a relationship - small, statistically detectable under its own model, and explaining under 2% of the variance at category level. Under 2% is a weak predictor, not an absent one, and the difference matters in both directions.

This page sets out what has been measured, by whom, with what interest, and what each figure is a percentage of. It is the review-platform part of the vendor-shortlist pillar. Where a figure's denominator does the work, the source-composition page owns the metric definitions and this page uses them without re-deriving them.

Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked platform documentation, quoted. Recommendation means a course of action that CoreAEX prescribes and no platform documents. Research findings are attributed in the prose with their sample, dates, engine and the outcome they measured, and are not tagged. Untagged text is ordinary explanation.

How Much Review-Platform Presence There Actually Is

Two measurements bracket the answer, and they are not the same measurement. One counts responses, the other counts articles, and both are single observations of a system that moves.

SE Ranking analysed 30,000 keywords producing 22,729 generated AI Overviews, on Google AI Overviews only, in the United States, from a snapshot dated December 1, 2025. It reports that out of all AI Overview responses, 34.5% cited at least one review platform - while those platforms represent just 8.5% of all links. A third of responses touch a review platform; review platforms are under a tenth of the links. Both are true, from the same runs on the same day, and quoting either without the other misstates the size of the phenomenon fourfold. SE Ranking's per-platform breakdown, which is the part most often quoted, consists of shares of review-platform links only - so a platform reported there at a large percentage still holds a small fraction of all links.

Citera measured something different on a much larger corpus. Analysing roughly 350,000 B2B SaaS articles across 10,382 keywords in 52 categories, querying ChatGPT and Claude with web search, Perplexity and Google AI Overviews in May 2026 on the US market in English, it reports that 19% of AI-cited articles are hosted on B2B review platform domains (G2, Capterra, TrustRadius), compared to 9% for non-cited articles. That is an over-representation of roughly two to one - and the authors attach two limits to it themselves, without prompting. The first: No domain authority control. We do not measure backlinks, Domain Rating, or domain traffic. Every comparison in this study is confounded by domain authority. The second: Each of the 10,382 keywords was queried once per AI engine.

Both limits bite. Review platforms are large, established, heavily-linked domains, so a comparison that cannot separate the platform from its domain strength cannot say which of the two the gap reflects. And the variation evidence shows single observations of these systems moving substantially between runs, which is a reason to treat any one-query-per-engine result as a wide-error-bar estimate rather than a rate.

The reading that survives both studies is narrow and still useful: in the measured datasets, review-platform domains are a recurring minority source class - present often enough to be worth understanding, and not the majority of links or of cited articles. That describes these prompt sets, these engines and these dates. Neither study sampled a universal population of B2B software recommendations, so it is not a source-share estimate for the category as a whole. On the wider picture of what else gets cited, the source-composition page has the full mix.

What Review Volume Predicts

One published regression addresses the question a SaaS team actually asks - whether collecting more reviews is associated with more citations - and it found a small association that was statistically detectable under its fitted model.

Kevin Indig, publishing on G2 Learn on October 23, 2025, analysed 30,000 AI citations and share-of-voice observations, drawn from Profound's tracking tool across 500 randomly selected G2 categories, using reviews from the preceding twelve months and citations from the preceding four weeks. The article does not name which specific AI engines Profound's tracking covered for this analysis. Categories were then excluded for fewer than 10 citations in the four-week window, a visibility score of zero, fewer than 100 approved reviews in the twelve months, or review counts that were significant outliers. The published page does not state the final regression sample after those exclusions, so 500 is the starting count rather than the analysed one. Two results:

Indig, G2 Learn, October 2025 - category-level regression; 500 categories sampled initially, final analysed sample after exclusions not stated
Relationship testedCoefficient95% confidence interval
G2 reviews and LLM citations0.0970.004 to 0.1910.009
G2 reviews and share of voice0.1130.016 to 0.2100.012

The author's own summary is a small but reliable relationship, and review volume explains less than 2% of the variation in citations and SoV. Read both halves of that. The reported 95% confidence intervals exclude zero, so the associations are statistically detectable under the study's fitted model and assumptions - which is not the same as being free of uncertainty, since sampling error, model specification, measurement dependence and omitted variables all remain, and the author lists each of them. The R² values are 0.009 and 0.012, so review volume accounts for a very small share of what varies between categories. What survives is a weak, model-dependent association rather than proof of an effect - and a weak predictor is still not the same as no predictor. The study licenses neither overstatement.

Four scope facts belong beside those numbers. The unit is the category, not the product - the author lists category-level aggregation among the study's limitations, noting it obscures differences between products inside a category, which is the level a vendor cares about. The design is cross-sectional, and the author states plainly that this means the estimates should be read as suggestive associations rather than causal effects. One platform is studied, so nothing here transfers to Capterra, TrustRadius, Gartner Peer Insights or any other. And the measurement depends on a third-party tracking tool, which the author also lists as a limitation. In all, seven limitations are stated on the page, including omitted-variable bias given the 98% of variance the model does not explain.

It is worth naming the publishing relationship directly. This regression appears on G2's own learning site, and it reports that G2 review volume explains under 2% of the variance in AI citations. That is a finding published against the publisher's commercial interest, which is a point in its favour and does not remove the interest. Definitions are given, which is more than most of the field offers: a citation is a site, G2 in this case, is cited in an LLM with a link back to it, and share of voice is the number of citations a site gets divided by the total available number of citations.

A separate G2 analysis, on its sales site, reports a large citation gap between paid and free profiles - and its own stratified results are the most useful part of it. This is a vendor publishing analysis of the product it sells, so the numbers need reading with the confounders in view rather than dismissing.

Published April 27, 2026, the analysis covers 84,623 products over 180 days: 7,532 paid profiles and 77,091 free listings, with citations tracked across ChatGPT, Perplexity, as well as Google's AI Overviews and AI Mode using a third-party tool named on the page. The aggregate figures are dramatic: the median paid G2 Profile earns 806 AI citations over 180 days. Free listings: 8. That's a 101x gap at the median, 15x at the mean.

Then the analysis partially adjusts for review volume, and reports the smaller number that results. Paid profiles carry more reviews than free ones, so the aggregate comparison mixes payment with review volume. Comparing paid and free within broad review-count bands instead - 0, 1-10, 11-50, 51-200, 201-500 and 500+ - the page states that the paid advantage at each review level is 2-7x, not 15x, and that a paid profile with 200 reviews will consistently beat a free listing with 200 reviews by about 2x. It also reports that reviews predict citations more than any other on-profile variable: r = 0.41 for paid products, r = 0.26 for free, and that profile completeness correlates with citations inside both tiers. The page's title, lead figure and adjusted figure are three different numbers - the H1 itself reads "2x," the subheading beneath it leads with the unadjusted 101x/15x gap, and the more conservative, band-adjusted 2-7x figure appears later in the body - which is unusual enough to note.

Bands are not matching. A band running from 51 to 200 reviews, or from 500 upward, still contains paid and free products that differ substantially in review count, and every confounder other than review volume remains untouched. This is useful stratification, and it is not a causal control.

What the analysis does not establish is the step from association to effect. Paid status is chosen, not assigned. Vendors who buy a profile differ from those who do not in ways the analysis does not measure - size, marketing budget, category maturity, how actively they solicit reviews, how much else they publish. Banding by review count narrows one of those differences and leaves the rest. And the page publishes analytical breakdowns without a reproducible data-collection method: it names its measurement tool, the four surfaces, the product counts and a 180-day window, while leaving the prompt set, the geography, a definition of what counts as a citation, the sampling process, the collection cadence and fixed start and end dates unstated. None of that makes the pattern false. It means the honest description is a substantial observed association in a vendor's own dataset, partially adjusted for one confounder, with no causal test.

The two G2 publications also sit at different levels and should not be read as one company contradicting itself. The regression is category-level over four weeks of citations; the paid-profile analysis is product-level over 180 days with a different comparison and a different denominator. Neither converts into the other, and no comparison between their coefficients is available.

What Google and Perplexity Document

No provider documentation located states that review platforms are preferred, weighted or treated as a distinct class in AI answers. Two documents are worth being precise about, because both are routinely stretched.

Google publishes a reviews system documentation page, and the sentence that matters most here is the one about what it leaves out. The system is designed to evaluate articles, blog posts, pages or similar first-party standalone content written with the purpose of providing a recommendation, giving an opinion, or providing analysis - and it does not evaluate third-party reviews, such as those posted by users in the reviews section of a product or services page. Documented That exclusion applies to user-submitted reviews on listing pages; it does not establish how Google treats every other element or page type on G2, Capterra or similar domains, which also carry vendor metadata, product descriptions, comparisons and platform-authored content. What the documentation covers is review articles; what it names as outside its scope is the user reviews in a listing's review section.

Two boundaries therefore apply at once, and both need stating. The documentation is classic Search guidance and makes no statement about AI Overviews, AI Mode or generative AI features - the only AI references on the page are navigation links to other documentation, checked across the page body and its sections on August 31, 2026. And within classic Search, it names the user reviews on a listing as outside its scope. So it documents neither special treatment of review-platform content nor any effect of it on AI answers.

That absence does not run in either direction. It does not mean review platforms are irrelevant to AI features, and it does not mean this system operates there. It means Google has documented a ranking system for review articles, excluded user reviews from it, and said nothing about AI answers, so any sentence connecting the three is an inference and should say so.

Of the provider documentation reviewed for this cluster, Perplexity is the only one publishing an inspectable source taxonomy. It names three labels: Government, the official website of a government organization; Academic, scientific sites; and Trusted, a broad category for sites that appear often enough in results to receive a label and that publish information within their own area of expertise. The stated purpose is reader orientation - labels help you see at a glance what kind of source you are reading. Documented There is no review-platform category and no commercial category. A review platform may of course receive the Trusted label like any frequently-appearing site; what does not exist is a documented tier for reviews.

One Widely-Quoted Figure This Page Does Not Use

One frequently-circulated review-and-AI statistic is absent here, and the reason is its metric rather than its access. It is worth explaining, because a reader who has seen the number deserves to know why it is missing.

A study commissioned by Trustpilot and conducted by Seer Interactive, tested in March 2026, analysed "804,491 AI responses" across "1,926 brands tracked across 8 industries," all US-based, on ChatGPT, Google Gemini, Perplexity and a fourth Google AI surface the report itself names inconsistently - "Google AI Mode" in its body and comparison chart, "Google AI Overviews" in its appendix. The full report is freely available as a PDF and was read in full on August 31, 2026. Its headline figures are given as rates: "1% citation rate for brands with no Trustpilot profile," rising to "75% Trustpilot citation rate for businesses that optimize their profile and collect a median of 81 reviews."

The report does not state what those rates count or what they are divided by. "Trustpilot citation rate" is ambiguous between the brand being cited and a Trustpilot URL being cited, and those are different measurements that would produce very different numbers. The accompanying press release resolves the ambiguity in a direction the report itself does not - "only 1% of answers from AI tools will cite a brand that has no Trustpilot profile" - prints 75.3% where the report prints 75%, and calls the fourth platform "Google AI Mode," which matches the report's own body wording but not its own appendix, where the same platform is called "Google AI Overviews." On confounding, the report states that it benchmarked cohorts using third-party SEO data (search volume, domain authority, backlinks and organic traffic) "to ensure fairness and isolate the effect of reviews from general SEO strength," but does not publish how that benchmarking entered the reported rates, and its cohorts are self-selected by profile activity rather than assigned.

The publisher also sells the remedy the finding recommends, which is worth stating but is not the reason for the exclusion. The reason is that a rate with no stated denominator cannot be compared with any other figure on this page, or used to size an effect. If Trustpilot publishes the definition, the figure becomes usable.

What Follows for a SaaS Team

The evidence supports a modest, unglamorous position, and it is better to say so than to manufacture a tactic. Three things follow, and only the first has any measurement behind it.

Treat review presence as table stakes, not as a lever. Review-platform domains appear in a meaningful minority of AI answers to commercial software questions, so an absent or badly out-of-date profile means one recurring source class carries no accurate description of your product. Fixing that is cheap. Expecting it to move a shortlist is not supported by anything located, and the one regression on volume found under 2% of variance explained at category level. Recommendation

Prioritise accuracy over volume. No study located measures whether review content is accurate, current or specific - every measurement here counts domains and citations, not what the pages say. That is a genuine gap, and in its absence the justification for keeping your listings correct is ordinary rather than algorithmic: buyers read them, your category page is often the first place a comparison happens, and a stale feature list or wrong pricing tier is a real cost whether or not any engine ever reads it. Recommend it on those grounds and do not dress it as a citation tactic. Recommendation

If you buy a profile, measure it as a change you made, not as a result you were promised. The paid-versus-free evidence is a vendor's own observational comparison with one confounder partially controlled. A team that upgrades and then reads a movement in its citations as proof has, at best, a single observation against a background of substantial run-to-run variation. The measurement page covers the panel design that makes such a before-and-after answerable at all: a fixed prompt set, repeated runs per engine per day, recorded conditions, and a comparison specified before the numbers were seen. Recommendation

What Is Not Known

The central question of this page has not been tested. No located study varies a review profile - adds it, removes it, changes its review count or its content - and measures what happens to a vendor recommendation. Every measurement here is observational, and three of the four are cross-sectional snapshots. An answer would need an intervention and a control, and nobody has published one for B2B software.

Nothing located measures review content, as opposed to review presence and volume. Whether detailed, recent, specific or critical reviews behave differently from thin ones is unaddressed by every study on this page, and it is the question a team with a mature review programme would most want answered.

All five research publications discussed here come from four commercially interested publishers. G2 published two of them - the regression and the paid-profile analysis; Trustpilot commissioned the excluded study; SE Ranking and Citera published one each, and both sell software and services in search and AI visibility. Named interest is not disqualifying, and the G2 regression in particular reports a result that does not flatter its publisher. It does mean none of this is independent measurement, and the independent multi-engine study that would settle the question does not currently exist.

And the platform side is documented only by omission. Google has published a reviews system for classic Search that explicitly excludes third-party user reviews and says nothing about AI features; Perplexity publishes three source labels and none of them is commercial or review-related. That is an absence of documentation, not evidence that review platforms are treated as ordinary sources.

Wondering whether your review programme is worth the investment for AI visibility?

The answer usually depends on what you would do differently either way - and that is a short conversation. Book a Session.

Sources

Sources: Platform documentation, quoted as read on August 31, 2026: Google Search Central, Google Search's Reviews System (last updated 2025-12-10 - classic Search ranking documentation covering first-party standalone review content, which states that it "does not evaluate third-party reviews, such as those posted by users in the reviews section of a product or services page"; its body makes no statement about AI Overviews, AI Mode or generative AI features, checked across the page body and its sections on the review date); Perplexity Help Center, Understanding source labels (last updated August 7, 2026). Vendor and practitioner research, each named as such: Kevin Indig, "Do More G2 Reviews Mean More AI Visibility? Insights from 30k Citations", published by G2 Learn, October 23, 2025 (30,000 AI citations and share-of-voice observations from Profound's tracking tool; twelve months of reviews against four weeks of citations; the specific engines Profound tracked for this analysis are not named in the article; category-level cross-sectional regression beginning with 500 randomly selected G2 categories and then excluding those with fewer than 10 citations in the window, a visibility score of zero, fewer than 100 approved reviews, or outlier review counts - the page does not state the final regression sample after those exclusions - the author states seven limitations including that the cross-sectional design means the estimates should be read as suggestive associations rather than causal effects, that category-level aggregation obscures product differences, that only one platform is studied, and that measurement depends on a third-party tracking tool); G2, "Why Paid G2 Profiles Earn 2x the AI Citations as Free Ones", published April 27, 2026 on G2's sales site (84,623 products over a 180-day window, 7,532 paid profiles and 77,091 free listings; ChatGPT, Perplexity, Google AI Overviews and AI Mode, tracked with a named third-party tool; the page's title, subheading and body each lead with a different figure - 2x, then 101x/15x, then the band-adjusted 2-7x; the page publishes analytical breakdowns - product counts, review-count bands, category comparisons and completeness buckets - but not a reproducible data-collection method, leaving the prompt set, geography, citation definition, sampling process, cadence and fixed start and end dates unstated - vendor analysis of the vendor's own paid product, observational, with review volume banded rather than matched and no other confounder addressed; figures printed here reproduced across three separate reads on August 31, 2026 after two earlier reads disagreed on one sentence, which is not printed); Citera, "An Analysis of 350,000 B2B SaaS Articles" (10,382 keywords in 52 categories; ChatGPT and Claude with web search, Perplexity and Google AI Overviews; May 2026, US market, English - the authors state that every comparison is confounded by domain authority and that each keyword was queried once per engine); and SE Ranking, "Despite 90% Traffic Loss, Review Platforms Top AI Overview Citations" (published January 29, 2026, last updated April 20, 2026; 30,000 keywords producing 22,729 generated AI Overviews, Google AI Overviews only, United States, data collected December 1, 2025; per-platform figures are shares of review-platform links only). Described but not relied on: the Trustpilot and Seer Interactive report "What AI says about you" (commissioned by Trustpilot, conducted by Seer Interactive, tested March 2026; 804,491 AI responses, 1,926 US brands across 8 industries, four engines; the full PDF is public and was read in full on August 31, 2026, including its appendix and its "About this study" section, where no limitations statement and no correlation-versus-causation statement was located in the Trustpilot-branded report itself; a companion write-up published separately by Seer Interactive does include one, noting that "it's possible that other factors like marketing spend and audience size influenced these rates"). No measurement cited here establishes that review presence, review volume or a paid profile causes a vendor to be recommended, and figures from different studies are not pooled or compared numerically anywhere on this page.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.