A wrong AI answer about your product is one observation, captured under conditions you mostly did not control and cannot fully see. It is an incident, not a measurement. What makes it useful later is how completely you recorded the conditions at the moment you found it, because you will not be able to reconstruct them afterwards and the same prompt may not return the same answer.

Two things happen on this page and one deliberately does not. You record the incident so it is reproducible, and you classify it so the next step is obvious. You do not explain it. A complete record tells you what was said, under what conditions, and how much it matters to a buyer. It does not tell you where the claim came from, and nothing you can capture from the interface will. This is the first step in a larger correction workflow, and the full ten-step sequence this record feeds into is set out separately.

Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked provider documentation, quoted. Recommendation means a field, category, threshold or workflow that CoreAEX prescribes and no provider documents - which is most of this page. Untagged text is ordinary explanation.

What to Record, and Why Each Field Matters

The record has two halves: the conditions the answer was produced under, and the claim itself. The conditions half is adapted from the condition schema used for prompt-panel work on the shortlist measurement page, restructured here for a single incident rather than a repeated panel run - an incident and a panel run should be comparable in spirit, even where the fields differ in detail. The claim half is specific to a factual error.

The fields below are a CoreAEX schema. No provider publishes one. Recommendation

The incident record - conditions
FieldWhy it is in the record
Prompt, verbatimThe exact characters you typed, not a tidied version. This is the input you controlled. It may not be identical to any downstream search query - OpenAI documents that ChatGPT search typically rewrites queries when it uses search partners
Engine and product surfaceChatGPT with search is a different surface from ChatGPT without it; AI Mode is a different surface from an AI Overview. "Google said" is not a record
Model or mode, where visibleRecord it where the interface exposes it. Where it does not, the observation date carries the scope instead
Account and login stateLogged in, logged out, which account, which plan tier
Search or browsing stateWhether the interface indicated it searched the web for this answer
Memory and personalisation stateOn or off, and whether the account has history the feature could draw on
Date, time and timezoneThe scope anchor for everything else in the record
Country and languageBoth as set in the interface and as the network would resolve them

Three of those fields exist because providers document behaviour that makes them matter. OpenAI's help documentation states that ChatGPT search typically rewrites your query into one or more targeted queries that it sends those providers, that ChatGPT may use your IP address to estimate your general location, such as your country, state, or city, and that if memory is enabled, ChatGPT may use relevant saved memories when rewriting a search query. Documented Those statements are about ChatGPT search specifically, and read at that scope they are enough to justify the fields: on that surface your prompt is an input to a rewriting step rather than necessarily the query that ran, and two people typing the same words from different locations or different accounts are not running the same test. Other engines document their own behaviour in their own terms, and the record captures the conditions rather than assuming any of them. Google separately documents that its Personal Intelligence memory experience in AI Mode is available in the US, in English - a reminder that personalisation features have their own availability boundaries, and that where an incident was captured is part of what it is.

The second half of the record is the claim.

The incident record - the claim
FieldWhy it is in the record
The disputed claim, quoted verbatim and isolatedOne sentence or clause, lifted out of the answer. Everything else in the response is out of scope for this incident
The full answer, as exported textContext, and the only way to spot a contradiction elsewhere in the same response
Displayed citations and displayed sources, as URLsThe starting point for tracing. Record them whether or not they look relevant
Whether the disputed claim carries a citation at allA separate field, because "cited" and "uncited" lead to different next steps - and neither settles where the claim came from
The correct value, and the authoritative URL for itIf you cannot point to where the correct value is published, that is itself a finding about your own evidence
Error type and severityThe two classifications below
Escalation required?Legal, security, compliance, contractual or safety. Recorded alongside the error type, not instead of it
How it surfacedA prospect on a call, a competitor's tweet, an internal check. This is not evidence of reach, but it is why anyone is looking
ScreenshotsSupplementary. A screenshot is not a substitute for exported text and a URL list

A field you could not capture is recorded as unavailable, not left blank and not guessed. "Model version not exposed" and "unknown" are different statements, and only one of them is true. An incident missing three condition fields is a weaker incident; it is still worth having, provided the gaps are visible to whoever reads it next.

Capturing the Answer Before It Moves

Capture is a sequence, and the common mistake is arguing with the model first. Recommendation Export the answer, then the source list, then the conditions - before you type anything else into the thread.

The reason is documented. Google's AI Mode help states that AI Mode might not always get personalization right. If AI Mode gets something wrong, you're always in control to respond with a follow-up to correct or clarify. Documented That is a real feature and a useful one, and it is worth being precise about what it does: it steers the response in front of you. It is not a correction to a source, and Google does not describe it as changing what anyone else is told. Once you send that follow-up, the conversation state has changed. Preserve the original answer before intervening, because any later response belongs to a different condition even if the first response is still visible in the thread.

Three practical points follow. Recommendation

Export text, not pixels. The disputed claim needs to be quotable and searchable later, and the source list needs to be a list of URLs rather than a picture of one. Screenshots are worth keeping for the interface state they show; they are not the record.

Record the run that produced nothing. Google documents that in some cases, AI Mode will provide a set of web links if there's not high enough confidence in the quality or helpfulness of an AI response. A response with no generated answer is a real state of the system, not a failed capture, and recording it as its own status keeps later denominators honest.

Try one reproduction, and record what happened either way. A fresh session, same prompt, same recorded conditions. If the claim appears again, you have observed recurrence under the recorded conditions. That still does not establish how often it appears, who else receives it, or whether buyers generally encounter it - those are panel and exposure questions, not conclusions from a second attempt. If it does not reappear, note that too, and do not conclude that the problem is gone.

Panel design, run counts and the observation window belong to a separate exercise - see how to test whether a correction changed an AI answer.

Six Error Types, Plus an Escalation Flag

The type decides who fixes it and what "fixed" would look like, so classify before routing. The six categories below are CoreAEX's, not any platform's, and their justification is operational: each one leads somewhere different. Recommendation

They describe what is wrong with a claim, not how much trouble it causes. Consequence is recorded separately, in the escalation flag below, because it cuts across all six - a misstated certification is both wrong and a compliance problem, and a category that tried to hold both would stop sorting cleanly.

The examples use an invented product, Acme Flow, with plans named Starter, Team and Scale. It is not a real company and none of these are observed incidents; they are there to make the boundaries between categories concrete.

The six error types, with an illustrative example of each. Acme Flow is invented
TypeWhat it meansIllustrative example
WrongStates a fact that is false against your authoritative sourceThe answer says single sign-on is included on the Team plan. It is available only on Scale
OutdatedStates a fact that was true and is no longerThe answer gives the Team plan at the price it carried before a change four months ago
Unsupported / unverifiedStates a checkable product or company claim that you have not yet substantiated and that the displayed citations do not support. It needs internal verification before it can be called true, wrong or outdatedThe answer says Acme Flow was founded by two former engineers from a named larger vendor. Nothing on any owned property says so, and no cited page contains it
IncompleteOmits a qualifier that changes the meaningThe answer says Acme Flow integrates with a named CRM, without the qualifier that the sync runs one way and is on Scale only
ConflatedMerges your product with another product, another tier, or another companyThe answer describes Acme Flow's features and attaches the pricing of a separate product with a similar name
ContradictoryContains two statements that cannot both be true, or contradicts its own citationThe answer says the Team plan has no API access, then two sentences later describes the Team plan's API rate limits

One of the six behaves differently from the rest and is worth separating deliberately. Unsupported is not a weaker form of wrong. It is the category for a claim you have not yet adjudicated, and its first action is a check of your own records rather than a correction. Until that check happens the claim is unverified, not false. Some claims in this category turn out to be true and merely unpublished, which changes the work entirely: the gap is on your side.

The Escalation Flag

Some incidents have to leave the marketing workflow before anyone drafts a correction, and the flag exists so that decision happens first. Record it alongside the error type - legal, security, compliance, contractual or safety - and route the incident to whoever owns that risk in your organisation. Recommendation

A flag is not a substitute for a type. An answer stating a certification you hold at a different level than the one claimed is wrong and carries a compliance flag. An answer describing a security control you retired last year is outdated and carries a security flag. Recording both keeps the content work legible once the escalation has been handled, and stops a flagged incident from disappearing out of the correction queue entirely.

The rest of this cluster is a content and evidence workflow. It is not built for the decision a flagged incident requires.

Severity, Judged Against a Buyer

Severity is not a property of the error. It is a property of the error in front of a particular buyer, which is why the record stores the buyer context alongside the score. Recommendation

  • Blocking. The error would disqualify you from consideration, or would survive into a contract. A missing capability the buyer has named as mandatory; a compliance or security claim; a price wrong by enough to reset a budget.
  • Distorting. The error changes the comparison without disqualifying you. A feature attributed to the wrong tier; an integration described without its limits.
  • Cosmetic. The error is visible and low-consequence for the decision. A founding year, a headcount, an out-of-date logo count.

The same claim moves between all three depending on who is reading. Single sign-on on the wrong tier is blocking for an enterprise security review, distorting for a mid-market comparison, and close to cosmetic for a five-person team that will never buy Scale. An incident scored without a named buyer context has been scored against an average buyer who does not exist.

Two habits keep the scale usable. Score against the buyer the prompt implies, where it implies one - a question about a 200-person company is not a question about a solo user. And treat severity as the input to sequencing rather than as a measure of how annoying the error is; the loudest error on a Monday morning is not automatically the most consequential one.

A Wrong Fact, or a Disagreement

An answer that recommends a competitor, ranks you low, or leaves you out entirely is not a factual error, and this cluster does not treat it as one. The test is whether a statement of fact is false, outdated, unsupported, incomplete, conflated or contradictory. Preference, ordering and inclusion are a different question with a different evidence base, and they belong to the vendor-shortlist cluster.

The boundary is easy to state and easy to blur under pressure, because both arrive looking the same - as a screenshot from someone in sales. Three cases sit right on the line.

"Best for enterprise teams" applied to a competitor is an evaluative claim, not a fact about your product. It is not an incident here, even when it is wrong-headed.

"Acme Flow does not support SAML" is a factual claim about your product, and it is an incident here even though it appears inside a recommendation and even though its practical effect is that you lost the comparison.

An omission is not a factual error, however costly. Being absent from an answer is the vendor-shortlist cluster's subject; there is nothing to correct on a page because nothing was said.

Where a single answer contains both - a false capability claim inside a recommendation for someone else - record the factual half as an incident here and take the rest to the other cluster. They are two pieces of work with two different definitions of success.

Why One Answer Is an Incident, Not a Rate

One captured answer establishes that this answer was produced once, under these conditions. It does not establish how often the claim appears, how many buyers have seen it, or that any other user would see it at all.

This matters because the pressure to generalise arrives with the screenshot. A single wrong answer forwarded by a founder becomes "ChatGPT is telling people our price is wrong" somewhere between the message and the meeting, and that sentence contains a rate that nobody measured.

Two things stand in the way of a single observation supporting a rate. The first is that the same prompt may not return the same answer between runs, and answers can vary across engines, accounts and locations - the variation page covers how much moves, at the scope of the studies that measured it. The second is that a rate needs a denominator: runs attempted, on a fixed panel, over a stated window. An incident has no denominator, and dividing by one is not a workaround.

What one incident does justify is the work. It is enough to start tracing, enough to check whether your own sources are right, and enough to open a correction. It is not enough to brief an executive on prevalence, to compare engines, or to claim a trend. If the question in the room is "how often does this happen," the honest answer is that the incident cannot say, and that measuring it is a separate exercise with a fixed panel and repeated runs.

That measurement - freezing a panel, capturing a baseline, and scoring runs at claim level - is covered on how to test whether a correction changed an AI answer.

Providers say something similar about their own output, in their own scope. OpenAI's documentation states that search results and citations can be incomplete, outdated, or incorrect, and Google's AI Mode help states that as with any early-stage AI product, AI Mode doesn't always get it right. Documented The same page continues directly: for example, in some cases AI Mode may misinterpret web content or miss context, as can happen with any automated system in Search. Those are acknowledgements that errors occur. Neither is a rate, and neither should be quoted as one.

What Happens Next, by Error Type

The record's purpose is to make the next step obvious without another meeting. The routing below is CoreAEX's, and it is deliberately coarse - the detail lives in the pages that own each step. Recommendation

Routing by error type
TypeFirst action
WrongConfirm the correct value against your authoritative source, then trace what the answer cited
OutdatedFind every place the superseded value is still published, starting with your own estate
Unsupported / unverifiedCheck your own records first, and reclassify once you know. If the claim is true and unpublished, the gap is yours; if it is not, search for where it is stated
IncompleteIdentify the missing qualifier and where it should have been visible next to the fact
ConflatedEstablish which product or company the wrong material belongs to before touching anything
ContradictoryCheck whether both halves exist somewhere in your own published sources

An incident carrying an escalation flag goes to legal, security or compliance before any of this, whatever its type. The routing above resumes once that owner has decided.

Every row leads to the same next question - what does the observable evidence around this answer actually show - and that is a separate piece of work with its own limits. Tracing the sources behind the claim is where that work starts; which candidate source to correct first, once several exist, is a separate decision this page does not cover.

What the Record Cannot Tell You

A complete incident record is a description of an output, and it stops there. Four limits are worth stating on the record itself, so that nobody downstream reads more into it than it holds.

It does not identify where the claim came from. The displayed citations are a place to start looking, not a verdict, and a claim can carry a citation that does not contain it. What can and cannot be established from the observable evidence is a separate exercise with its own evidence tiers, covered in how to trace the sources behind a wrong claim.

It does not distinguish a retrieved error from a generated one. Whether the claim came from a page the engine consulted, from something the model held already, or from the two being combined is not visible from the interface, and no capture procedure will make it visible.

It does not tell you what other people see. Providers document that answers can vary with location, account state and memory. Your capture describes your session.

And it does not predict tomorrow. An incident is a dated observation. Whether the claim persists, disappears on its own, or returns after you fix something is a question for a panel over a stated window, not for a record of one answer.

Want the incident record as a template your team can actually fill in?

It is one page of fields, and most of the value is in the three that people skip. Book a Session.

Sources

Sources: Provider documentation, quoted as read on September 1, 2026. OpenAI Help Center, ChatGPT search - the query-rewriting, IP-location, memory and "incomplete, outdated, or incorrect" statements quoted above. This article displays only a relative update date rather than a fixed publication or revision date, so it is cited to the review date. Google Search Help, Get AI-powered responses with AI Mode in Google Search - which carries "AI Mode doesn't always get it right," the low-confidence web-links behaviour, and the follow-up sentence quoted in the capture section; no last-updated date is displayed on this page, so it too is cited to the review date. A fourth passage on that page, about misinterpreting web content, was initially found inconsistent across two reads; a repeat check on the review date returned it identically twice, confirming it as stable and complete, and it is quoted directly in the body above rather than paraphrased. Google Search Help, Personal Intelligence: How AI Mode in Search personalizes responses for you - the memory experience is stated as "available in the US, in English"; no last-updated date is displayed. Provider help pages change without notice and these three are re-checked before publication and quarterly thereafter. What none of these documents: how often any engine states a product fact incorrectly, which kinds of error are most common, or whether a captured answer would be reproduced for another user. No figure on this page is a rate, because no rate was located. A deliberate search for an independent, methodologically transparent audit of AI-answer accuracy about SaaS product, pricing or feature information returned news-accuracy audits and vendor-commissioned studies whose figures could not be verified at the publishing vendor's own site; no product-specific audit meeting those criteria was located, and the categories and severity bands on this page are therefore presented as an operational classification rather than as a measured distribution. The incident record, the six error types, the escalation flag, the severity bands and the routing table are CoreAEX's own and are not documented by any provider or platform. All examples use an invented product and are illustrative; they are not observed incidents and no client is described.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.