Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked provider documentation, quoted. Recommendation means a procedure or classification that CoreAEX prescribes and no provider documents. Research findings are attributed in the prose with their sample, engines, dates and the outcome they measured, and are not tagged. Untagged text is ordinary explanation.

Tracing a wrong claim means assembling the evidence you can actually observe - the displayed citations, the search results, the passages that match - into a plausible source chain, and being explicit about which links in that chain you have demonstrated and which you have inferred. It does not mean finding out where the claim came from, because from outside the system that determination is not available.

A citation is a place to look, not a verdict. The rest of this page is a procedure for looking, and a vocabulary for writing down what you found without overstating it. The single most useful habit is the one that sounds least satisfying: labelling a finding "plausible contributor" when that is what the evidence supports, and resisting the sentence everyone in the room wants, which is that you have established where the claim came from. Tracing sits partway through a longer workflow, and where it fits alongside capturing, correcting and re-testing is mapped out on its own page.

Three Tiers: Demonstrated, Inferred, Unknown

Sort every finding into one of three tiers before it goes into a write-up, and keep the label attached to it afterwards. Recommendation The tiers are not a formality - they are what stops a trace from being read as an explanation.

The three evidence tiers
TierWhat belongs in itTest
DemonstratedThe answer displayed these URLs. This page contains this passage. This page displays this publication or update date. Your own pricing page states this valueYou can show it to someone else and they will see the same thing
InferredThis third-party listing is a plausible contributor to the wrong value. The claim most likely reflects the tier structure as it stood before the March changeA reading consistent with the evidence, which the evidence does not isolate from other readings
UnknownWhat the engine retrieved but did not display. Whether an uncited page contributed. Whether the claim came from retrieved text or from the model. Everything inside the provider's pipelineNo procedure available to you would settle it

The third tier is the one that gets quietly emptied. A trace with a populated unknown column reads as incomplete work, so there is steady pressure to promote its contents into tier two and tier two into tier one. Resisting that is most of the discipline here, and the write-up section below is built to make the unknowns survive compression.

Reading the Citations, and Testing Them Against the Claim

Start from the captured incident, not from a new run. The captured incident already holds the disputed claim quoted in isolation, the displayed sources as URLs, and whether the claim carried a citation at all. Reopening the saved conversation lets you inspect the recorded answer; regenerating it, resubmitting the prompt or starting a fresh chat creates a new observation under potentially different conditions rather than a closer look at the original one.

Then run the test that the rest of this page depends on. Recommendation For the disputed claim specifically - not the answer as a whole - open each cited source and look for the claim in it. Search the page for the wrong value. Three outcomes, and they lead in different directions.

  • A cited source contains the claim. You have a demonstrated fact: this URL carries this wrong value. That makes it a correction target and a plausible contributor. It does not make it the origin.
  • The claim is cited, but the cited source does not contain it. Citation-support failures have been documented in the audits below. The visible citation therefore identifies material to inspect; it does not establish where this sentence came from, and the search in the next section becomes the main line of work.
  • The claim carries no citation. Nothing follows from this about whether a web source was involved. Absence of a visible citation is not evidence that no source contributed.

One more thing constrains what the citations can tell you: the query the engine ran was probably not the prompt you typed. OpenAI's documentation states that ChatGPT search typically rewrites your query into one or more targeted queries that it sends those providers. Documented That is one provider describing one surface, and it is enough to make the point - the sources you are looking at were selected against something you did not see and cannot reconstruct.

What the Citation-Support Research Shows, and What It Does Not

The reason claim-to-citation matching is the first step rather than a formality is that citations in generative search have been measured, repeatedly, not supporting the sentences attached to them.

The peer-reviewed human evaluation used here is now three years old, and its age matters. Nelson Liu, Tianyi Zhang and Percy Liang, in Findings of the Association for Computational Linguistics: EMNLP 2023, ran a human evaluation across 1,450 queries on four generative search engines - Bing Chat, NeevaAI, perplexity.ai and YouChat - and report that "on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence." Two of those four systems no longer exist in the form tested, so these figures describe what was measured then, not a current rate for any live product. The authors also draw a boundary worth keeping: "verifiability is not factuality" - whether a statement can be checked against its cited source is a different question from whether it is true.

A 2026 preprint measures something closer to the operational problem, and its result is the one to hold onto. Hailey Onweller, Elias Lumer, Austin Huber, Pia Ramchandani, Vamse Kumar Subbiah and Corey Feld benchmarked 14 closed- and open-source models on three separate dimensions - whether the link works, whether the content is topically relevant, and whether it factually supports the claim - and report that "even the strongest frontier models maintain link validity above 94% and relevance above 80%, yet achieve only 39-77% factual accuracy." Read narrowly, that is the practical lesson of this page: a citation can be live, on-topic, and still not contain what the sentence says. The paper is a preprint, not peer-reviewed, and its own limitations section states that the LLM-as-a-judge approach used for the relevance and fact-check dimensions "may retain biases inherent to the judge model, including position bias and self-enhancement effects," and that cited URLs are temporally unstable - pages accessible during evaluation may later disappear.

A separate preprint shows that the sources an answer engine displays and the sources it cites inline are not necessarily the same visible set. Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Yixin Mao and Chien-Sheng Wu ran a study with 21 participants across three answer engines - YouChat, Bing Copilot and Perplexity AI - and report the gap as counts rather than rates: a mean of 4.31 displayed sources against a mean of 3.0 cited, with Perplexity displaying the most (mean 5.00) while citing the fewest (mean 2.58), and YouChat citing all of what it displayed (mean 3.57).

That is an interface-level gap, and it is easy to over-read. It establishes that a source may be shown without being cited. It does not establish that a displayed-but-uncited source was consulted by the model, entered its generation context, or shaped any sentence - the study counted what the interface exposed, not what the provider retrieved internally, and no study located makes that link. When you find yourself reasoning that the real source must be something the engine looked at and did not show you, that reasoning is tier three.

One boundary, and the scope differs by study. Liu and colleagues audit citation support in four generative search engines; Venkit and colleagues compare displayed source lists with inline citations in three answer engines; Onweller and colleagues evaluate source attribution in model-generated deep-research reports. None of the three measures how often AI answers get product, pricing or feature facts wrong, and we located no independent, methodologically transparent audit that does. They tell you why to check the citation. They do not tell you how likely your incident was.

When the citations do not contain the claim, the next move is to look for the claim itself on the open web, and the trick is to search for the wrong value rather than the right one. Recommendation

Four searches cover the cheap options.

  • Exact-phrase search on the distinctive wording. Take the most unusual six or eight words of the claim as the answer stated them and quote them. Distinctive phrasing travels; generic phrasing returns everything.
  • The wrong value with the product name. The superseded price, the feature attributed to the wrong tier, the integration described the wrong way - paired with the product name and nothing else.
  • Site-restricted searches against the properties most likely to carry a stale version: review platforms, directories, comparison pages, your own legacy URLs.
  • Your own estate, honestly. Run the same search restricted to your domain before assuming the problem is external. An old campaign page or a documentation section that never got updated is a common and unwelcome result.

A match gives you a plausible contributor and nothing stronger. A page containing the wrong value in wording close to the answer's is a tier-two finding: it is consistent with that page having contributed, and it does not exclude the several other pages that also carry the value, or the possibility that neither was involved.

No match is a real result and changes the plan. If a deliberate search across the formulations above locates no source stating the claim, record that - naming the searches you ran, so the absence is a statement about a defined search rather than about the web. Practically it means there may be nothing external to correct, which redirects the work toward your own sources and toward the provider feedback routes.

Canonical, Syndicated and Superseded Copies

Once you find the value somewhere, expect to find it in several places, and separate the question of which copy matters from the question of which copy was involved.

Product facts syndicate. A directory entry gets scraped into an aggregator, a listicle quotes a review profile, a partner page copies your own old wording. Two consequences. Recommendation

For the correction, find the copy nearest the fact's owner - which is the next page's subject, not this one's. For the trace, note that a value appearing on nine URLs makes any single-URL attribution weaker, not stronger.

For the timeline, check what each page said when the incident was captured, not only what it says now. A page you read today may have been corrected since, and a page that is right now may have been wrong then. Where an archived copy exists, it is demonstrated evidence of the earlier state; where none exists, the current read is evidence about today and the earlier state is unknown. A source corrected after your incident can still be the plausible contributor to it, and a trace that only looks at present-day pages will miss that entirely. Recrawl, re-retrieval and what providers do and do not document about any of it is covered on its own page.

Retrieval or Model Knowledge: The Question That Stays Open

Everyone asks whether the wrong claim came from a page the engine read or from something the model already held. From outside the system, that is not distinguishable, and no procedure on this page will settle it.

The research literature at least supplies clean vocabulary for the possibilities. Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang and Wei Xu, in a survey at EMNLP 2024, name three categories of knowledge conflict: context-memory conflict, between retrieved content and what the model holds; inter-context conflict, between two retrieved sources; and intra-memory conflict, within the model itself. Those are useful labels for describing what you cannot rule out. They are a survey's taxonomy, not a finding about your incident.

One controlled result explains why finding a matching passage is still worth the effort. Kevin Wu, Eric Wu and James Zou, at NeurIPS 2024 in the Datasets and Benchmarks Track, built a dataset of over 1,200 questions across six domains, took real source documents and deliberately replaced the answers in them, and benchmarked six models including GPT-4o. They report that models are "susceptible to adopting incorrect retrieved content, overriding their own correct prior knowledge over 60% of the time." That figure comes from a dataset built to contain errors: the authors state that it "contains an enriched rate of contextual errors, so the reported metrics are not meant to represent bias rates in the wild." So it does not tell you how often this happens in production, and it is not a claim about your incident. What it supports is narrower and still useful: a wrong page sitting in the retrieved context is capable of overriding a correct prior, which is a reason to take a matching third-party passage seriously rather than dismissing it because the model "should know better."

The honest write-up position is that the stage is unknown. Repeated measurement does not close this either - a correction panel can tell you what changed and how consistently, and it still cannot locate where in the pipeline anything happened.

Writing the Trace So the Unknowns Survive

A trace write-up has one job beyond recording what you found: it has to stay accurate when someone summarises it in a sentence. That sentence is where traces go wrong, because the compressed version of "a third-party listing carries a matching passage" collapses into "we found the source."

Five fields, in this order. Recommendation

  • The disputed claim, quoted as the answer stated it.
  • The displayed sources, as URLs, and whether the claim carried a citation.
  • The claim-to-citation result - which of the three outcomes in the citations section applies.
  • Candidate contributors, each with its tier label and the evidence for it.
  • What could not be established, stated positively rather than left blank.

And a vocabulary rule that does more work than the rest of the page combined. Recommendation

Wording for a trace write-up
Do not writeWrite
"The source of the error is X""X contains a passage matching the wrong claim"
"The AI pulled this from X""The answer cited X"
"X is why the answer is wrong""X is a plausible contributor; we could not establish that it was used"
"There was no source for this""We located no page stating this claim, across these searches"
"The model made it up""We could not determine whether the claim came from retrieved content or from the model"

These are not softenings. Each right-hand phrasing is a claim you can defend if someone opens your evidence; each left-hand one asserts a link to the provider's internal process that you did not observe. The right-hand column is also more useful downstream, because it tells the next person what was checked rather than handing them a conclusion they cannot audit.

With the trace written, the question becomes which of the candidate sources to correct and in what order - a decision about ownership and authority rather than about evidence.

Want a trace that survives the executive summary?

Most of the work is in the second and third columns - separating what you can show from what you are reading into it. Book a Session.

Sources

Sources: Provider documentation, quoted as read on September 1, 2026: OpenAI Help Center, ChatGPT search - the query-rewriting statement quoted above; the article displays only a relative update date rather than a fixed publication or revision date, so it is cited to the review date. Peer-reviewed research. Nelson Liu, Tianyi Zhang and Percy Liang, "Evaluating Verifiability in Generative Search Engines", Findings of the Association for Computational Linguistics: EMNLP 2023 - human evaluation, 1,450 queries, four generative search engines (Bing Chat, NeevaAI, perplexity.ai, YouChat), the source of the 51.5% and 74.5% figures quoted above; two of the four systems no longer exist in the form tested, and the authors state that "verifiability is not factuality". Kevin Wu, Eric Wu and James Zou, "ClashEval", Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track - over 1,200 questions across six domains, six models, real documents with answers deliberately replaced; the "over 60%" figure comes from a dataset the authors state "contains an enriched rate of contextual errors, so the reported metrics are not meant to represent bias rates in the wild," and is quoted on this page only with that sentence attached. Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang and Wei Xu, "Knowledge Conflicts for LLMs: A Survey", EMNLP 2024 - cited for its three-category taxonomy only, not for a finding. Preprints, labelled as such and not peer-reviewed. Hailey Onweller, Elias Lumer, Austin Huber, Pia Ramchandani, Vamse Kumar Subbiah and Corey Feld, "Cited but Not Verified", arXiv 2605.06635, submitted 7 May 2026 - 14 models across three evaluation dimensions, the source of the link-validity, relevance and factual-accuracy figures; its limitations section is quoted above on judge-model bias and URL instability. Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Yixin Mao and Chien-Sheng Wu, "Search Engines in an AI Era", arXiv 2410.22349 - 21 participants across YouChat, Bing Copilot and Perplexity AI; its displayed-versus-cited figures are means of counts, not percentages, and are reported here as the paper reports them. Scope differs by source, and the five are not interchangeable. Liu and colleagues audit citation support in four 2023-era generative search engines. Venkit and colleagues compare displayed source lists with inline citations in three answer engines - an interface-level measurement, not an observation of what a provider retrieved internally. Onweller and colleagues evaluate source attribution in model-generated deep-research reports. ClashEval is a controlled benchmark of conflict between a model's prior and deliberately altered contextual evidence, not an audit of live systems. Xu and colleagues is a survey, cited only for its taxonomy. None measures the accuracy of AI answers about product, pricing or feature information, and we located no independent, methodologically transparent audit that does. Nothing on this page establishes where any individual claim originated; the tier labels exist precisely because that determination is not available from outside the system.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.