There's no single number that answers "are we getting cited." Mentions, citations, impressions, and share of voice measure genuinely different things, and the tools that report them don't all define them the same way - a common measurement mistake in this space is treating those terms as interchangeable, or assuming a rising citation count automatically means rising traffic or revenue. It usually doesn't move in lockstep. This page works through what's actually measurable per engine, where the available tools' own definitions diverge, why citation volume and referral traffic are separate questions, and what a recent academic survey of this exact field concluded about the limits of what any of this can currently prove.

Start With a Metric Dictionary - Don't Assume the Terms Match Across Tools

Before comparing any two numbers, check whether they're measuring the same thing. Semrush defines its AI-visibility metrics with reasonable precision: "Mentions" is the number of prompts in which a brand appears in AI responses; "Citations" is the number of AI responses that cite the brand's domain as a source; "Share of Voice" is the brand's percentage of mentions relative to a defined competitor set; "Cited Pages" are the specific pages referenced. Semrush's documentation counts these as separate metrics rather than defining one in terms of the other, and that separation matters in practice: a response can mention a brand by name without citing any source at all, and - depending on how a given tool scopes "citation" - a response can cite a domain without the brand name appearing in the visible answer text. Don't assume a rising citation count implies a proportionally rising mention count, or the reverse, without checking both numbers directly. Ahrefs' own Brand Radar methodology page is notably more candid about the limits of its own numbers: it doesn't give "mention" or "citation" a formal definition beyond describing them operationally as string and link matches within a stored corpus of AI responses, and it states directly that "Estimated Impressions" is "a modeling choice, not a measured relationship" - Ahrefs is explicit that it doesn't claim a validated link between Google search volume and how often a given query is actually asked inside an AI tool. That's a genuinely useful disclosure, not a weakness to paper over: it means a raw "impressions" number from any vendor tool is a model output, not a census count, and should be read that way.

What You Can Measure Directly, and What's a Sampled Panel

Two different kinds of numbers get reported under "AI measurement," and they're not interchangeable. The first kind is first-party, observed data - a platform reporting what actually happened on its own system. Google Search Console has a dedicated generative AI performance report, covering impressions in AI Overviews and AI Mode, broken down by page, device, and country. It has real limits worth knowing before relying on it: it doesn't show which specific passage was cited or any competitor data, it requires a minimum impression volume before data appears at all, and Google says the report is still rolling out rather than universally available - but the impressions it does report are counted, not modeled. For ChatGPT specifically, OpenAI's own publisher documentation confirms a concrete, checkable mechanism: "ChatGPT automatically includes the UTM parameter utm_source=chatgpt.com in referral URLs," which means ChatGPT-driven traffic is directly visible in standard analytics tools without needing a third-party AI-tracking product at all.

The second kind is a sampled, modeled panel: a tool runs a defined set of prompts against each engine on a schedule and reports what came back, which is a genuinely different measurement than a census of every real user query. Ahrefs Brand Radar and Semrush's AI visibility tracking both work this way for cross-engine coverage. Their per-engine coverage isn't uniform, either - per Ahrefs' own Brand Radar help documentation, the platform currently tracks seven groupings (Google AI Overviews and AI Mode counted together, plus ChatGPT, Perplexity, Gemini, Copilot, Grok, and Claude), and two of those carry real caveats worth checking before trusting a cross-engine comparison: Ahrefs has stated it is "temporarily unable to gather new data" for Grok following a platform policy change, and Claude is tracked only via API on custom prompts, not as part of the standard automated panel. Confirm a tool's current per-engine coverage and any collection gaps before using it to argue a brand is under- or over-performing on a specific engine - that documentation changes as platforms change their terms.

Why More Citations Doesn't Automatically Mean More Traffic

A rising citation count doesn't reliably translate into a rising click count. Pew Research tracked the actual browsing behavior of 900 US adults across 68,879 real Google searches in March 2025: when an AI summary appeared in the results (12,593 of those searches), users clicked a traditional search result link in 8% of visits, compared to 15% of visits when no AI summary appeared - roughly half the click rate. Clicking a link inside the AI summary itself happened in just 1% of visits to a search results page that displayed one. That 1% figure is a share of visits to AI-summary result pages, not a per-citation click-through rate - an AI summary can contain several cited links, and the 1% describes how often any of them got clicked, not the odds that one specific citation gets clicked. The same research team's fuller published analysis of this dataset (Chapekis, Lieb, Shah & Smith, arXiv, August 2026) reports the identical 1%-of-visits figure and adds that the association between AI-summary presence and fewer outbound clicks holds up under a mixed-effects logistic regression controlling for panelist-level random effects and the query attributes that make an AI summary more likely to appear in the first place - which strengthens confidence that this is a real association, though it's still an observational one, not a randomized comparison. For a more direct causal test, a separate preregistered field experiment (Wang, Gleason, Bart, Wilson & Metaxa, arXiv preprint, August 2026, N=1,100) randomly assigned users to search with or without AI Overviews and AI Mode, and found that removing those features increased click-through to publisher sites - a design that can support a causal reading in a way an observational panel can't, though as a preprint it hasn't yet gone through peer review. Together, these are the concrete reason citation rate, mention rate, and referral traffic need to be tracked as separate metrics rather than assumed to move together. One scope note on the original Pew figures: this is a Google-specific, US-panel, March-2025 snapshot - Pew notes it couldn't reliably identify AI summaries on other search engines, and both AI-summary behavior and interfaces have likely changed since. Treat the direction (more AI-answer exposure, fewer outbound clicks) as the better-supported takeaway than the exact percentages a year or more later.

Track Citation Accuracy, Not Just Citation Volume

A citation count alone doesn't tell you whether an AI engine is representing a brand or its content correctly. A Tow Center for Digital Journalism study ran 1,600 tests across eight AI search tools, asking each to identify the correct article, publisher, date, and URL behind a direct excerpt from twenty news publishers. Collectively, more than 60% of queries got an incorrect answer, with error rates ranging from 37% (Perplexity) to 94% (Grok 3); one tool misattributed the source 115 times out of 200. Confident-sounding wrong answers were common: ChatGPT was wrong on 134 of its 200 test responses, and across all 200 responses - right and wrong combined - it signaled a lack of confidence only 15 times, and never once declined to answer outright. The study's own scope is specific and worth preserving: it tested news-article retrieval only, ran each excerpt through each tool once rather than repeatedly, and the authors explicitly caution against extrapolating the exact figures to every model or every kind of query. The practical implication still generalizes: when checking how an AI engine represents a brand, look for misattribution and factual accuracy, not only whether the brand's domain shows up somewhere in the response.

What "Share of Voice" Actually Means

"Share of voice" isn't a single standardized formula - it's whatever a given tool defines it as, and the underlying calculation differs by vendor. Semrush, for instance, defines its Share of Voice specifically as a brand's mentions as a percentage of total mentions across a defined competitor set - a mentions-based figure, not a citations-based one. A different tool could build the same-sounding metric around citations instead of mentions, or weight it differently, and produce a different number for the same brand over the same period without either figure being wrong. Whichever tool a report comes from, check that specific tool's definition before treating "share of voice" as a portable, cross-tool number - and before using a share-of-voice figure to argue a campaign worked, confirm the competitor list and prompt panel didn't also change between the before and after measurement.

A Practical Audit Checklist for AI-Citation Readiness

In rough order: first, confirm the technical basics - crawlability, indexing, and that AI crawlers aren't blocked (covered in full in CoreAEX's crawlability pillar). Second, measure the activation rate for the target query set - the percentage of a defined, fixed set of queries that trigger a generative answer at all, checked across repeated runs rather than a single pass, since some AI answers are inconsistent run to run. Activation rate is worth tracking as its own number precisely because it can move independently of anything a single page does; a low rate on a given query set is information, not necessarily a sign that on-page work is misdirected. Third, check factual currency and completeness on the pages that matter most (covered in depth on the repurposing page). Fourth, pull current mention and citation data from at least one cross-engine tool, reading its specific metric definitions first rather than assuming they match another tool's. Fifth, spot-check accuracy, not just volume - does the AI engine represent the brand and its facts correctly when it does cite or mention them. Sixth, track referral traffic and mentions as separate line items, since one rising doesn't imply the other is. Seventh, keep the measurement protocol itself reproducible: use a frozen prompt set rather than a shifting one, run each prompt more than once before treating a single response as representative, record both the citation rate and the no-citation/non-activation rate rather than just the former, and canonicalize URLs before comparing cited-page counts across periods - otherwise apparent trend changes can just be measurement-method noise.

The Honest Limit of All of This

A 2026 critical survey reviewing 45 studies in this field over roughly three years reaches a conclusion worth stating plainly: "already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior." That survey is a single-author arXiv preprint, not a peer-reviewed publication, and its conclusion is explicitly bounded to the 45 studies it reviewed (November 2023-July 2026) - it's a critical read of the existing evidence base, not an independent experiment of its own, and a stronger or more current study could narrow the gap it describes. With that caveat, the practical reading still holds: the measurement tools above can tell you what's true right now, on the engines and prompt panels being tracked, and controlled experiments - including the preregistered field experiment cited above - can show a real effect within their own tested conditions. None of that currently proves that a specific on-page change reliably, durably increases citations across every engine over time. Measurement here means tracking observed state accurately, not proving causation - treat any dashboard's month-over-month trend as a description of what happened, not a settled explanation of why.

← Back to the AEO fundamentals pillar
How to repurpose existing content for AEO →
How AI crawlers actually access your site →


Sources: Mention, citation, share-of-voice, and cited-pages definitions are from Semrush's AI Visibility Metrics documentation. Brand Radar's operational definitions and its own "modeling choice, not a measured relationship" disclosure on Estimated Impressions are from Ahrefs' Brand Radar Methodology page; current per-platform coverage and collection status (seven groupings; Grok collection temporarily paused; Claude tracked via API on custom prompts only) are from Ahrefs' Brand Radar help documentation, which should be checked again before quoting exact coverage, since platform terms and collection status change. Google Search Console's generative AI performance report scope and limitations are from Google's own Search Console Help documentation. ChatGPT's automatic UTM tagging is from OpenAI's Publishers and Developers FAQ. The click-rate figures (8% vs. 15% for traditional results with and without an AI summary; 1% of visits to AI-summary pages for clicks inside the summary itself) are from Pew Research Center's study (900 US KnowledgePanel members, 68,879 Google searches, March 2025) and the same research team's fuller analysis, Chapekis, Lieb, Shah & Smith (arXiv, August 2026), which adds the mixed-effects regression controls - both a Google-specific, US-panel snapshot; Pew notes it couldn't reliably identify AI summaries on other search engines. The causal field-experiment finding that removing AI Overviews/AI Mode increased publisher click-through is from Wang, Gleason, Bart, Wilson & Metaxa (arXiv preprint, August 2026), a preregistered experiment (N=1,100) not yet peer-reviewed. The citation-accuracy figures (over 60% of queries answered incorrectly collectively; 37%-94% error rates by tool; 115/200 misattributions; ChatGPT wrong on 134/200 and signaling uncertainty on 15/200 overall) are from the Tow Center for Digital Journalism's study (1,600 tests, 8 AI search tools, 20 news publishers) - scoped to news-article retrieval specifically, tested once per excerpt, not claimed to generalize to every model or query type. The measurement field's stated causal limits ("no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior") are from Olivier Martinez's critical survey - a single-author arXiv preprint, not peer-reviewed, reviewing 45 studies (November 2023-July 2026); its conclusion is bounded to that reviewed set.

About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.