Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked provider documentation, quoted. Recommendation means a protocol, threshold, scoring rule or reporting format that CoreAEX prescribes and no provider documents - which is nearly all of this page. Research findings are attributed in the prose with their sample, engines and the outcome they measured, and are not tagged. Untagged text is ordinary explanation.

Testing whether a correction changed an AI answer is a sampling problem, and it is a sampling problem without a control condition. A frozen prompt panel, a captured baseline, repeated runs and claim-level scoring will tell you what changed and how consistently it changed. None of it will tell you why.

That gap is not a gap in your rigour. We located no published peer-reviewed paper or preprint that intervenes on a web source and measures whether an AI answer subsequently changed. A controlled version may exist privately or remain unpublished; none was available to support this protocol on the review date. So the design below is not a weaker version of a rigorous one that exists elsewhere in the literature - it is what is available, and it does not establish cause. What a good protocol buys is a defensible description of what happened and an honest account of what remains unknown, which is enough to decide what to do next and not enough to claim a mechanism. This panel is the re-test step of a larger correction workflow, and the full sequence it sits inside, from capture through reporting, is covered on its own page.

What You Are Measuring, and What You Are Not

"Did the correction work" is not one question, and collapsing it into one is the mistake that makes a result unreportable. A correction passes through a sequence of separate states, each with its own evidence, and they do not move together. Recommendation

The outcome ladder. Each row is recorded separately
StageOutcomeWhat establishes it
RequestCorrection requestedA dated submission through a named route
Request acknowledgedThe publisher or platform confirmed receipt
Request accepted or refusedThe publisher or platform said so
SourceSource correctedThe target URL now carries the correct value
Correction date verifiedWhen the corrected version became publicly reachable, and how you established that
Post-correction observationCrawler fetch observedVerified logs show a named crawler requested the URL after the correction date. This establishes that request and nothing else - not that the answering system consulted the page
Corrected URL displayed or citedThe corrected URL appears in the answer interface. This does not establish that the page was retrieved by the answering system, or used for the disputed claim
AnswerAnswer correctedThe disputed claim is stated correctly
AttributionCause or attributionNot observable from outside the provider. The row exists so that nobody records one of the rows above in its place

These need not move together, which is why they are separate rows rather than a progress bar. A crawler fetch, a displayed citation and a corrected answer can diverge in any direction, though some states carry logical dependencies - a corrected-source citation presupposes the source was corrected. An answer can start stating the correct value with no observable fetch or citation of anything you changed. Recording only the last row throws away the evidence that would let you tell those situations apart.

One gate matters more than the rest: only incidents that reached a verified source correction belong in the answer analysis. A request that was refused, ignored or still open is a real and reportable finding - it is the honest answer to "what happens when you ask" - but it is not a correction test, because nothing was corrected. Report those on their own denominator of requests submitted, and keep them out of the numerator of anything else.

Two further outcomes sit outside this page. Whether your brand is mentioned, and whether a recommendation changed, are questions about inclusion rather than accuracy, and the shortlist measurement page owns them. Record them if you like; do not report them as evidence that a factual correction landed.

Building and Freezing the Panel

The panel is a fixed set of prompts that you will run unchanged for the length of the study. Freezing it is what makes the before and after comparable, and editing it midway breaks before-and-after comparability, which means the original test is no longer the one being run. Recommendation

Seed it from the incident. The incident record already contains the prompt that produced the wrong claim, the engine and surface it appeared on, and the full condition schema - those are the starting parameters, not something to reconstruct later.

Use one primary prompt per claim, and buy reliability with repeated runs rather than with prompt variants. The temptation is to write six phrasings of the same question on the theory that breadth is rigour. It is not: six phrasings run once each give you six single observations of six different things, while one phrasing run six times gives you a distribution of one thing. Variants answer a different question - whether the claim appears across phrasings - and if you want that question, treat each variant as its own frozen prompt with its own runs and its own denominator.

Four practical rules hold the panel together. Recommendation

  • Version it and date it. If a prompt has to change, that is a new panel version and a new series. Say so in the write-up rather than splicing the two.
  • Name every engine and product surface, and report each separately. Never pool results across engines. Engines are not replicates of each other.
  • Name what you excluded. An engine you left out is stated as excluded, not silently absent - otherwise your write-up implies coverage you do not have.
  • Record the full condition schema on every run. Account state, location, language, memory setting, date and time. A run whose conditions were not recorded cannot be compared with anything.

The Baseline, Captured Before Anything Changes

A correction test needs a before, and the before has to be a distribution rather than the original incident. One captured answer is an incident. Comparing it against thirty post-correction runs compares a single observation with a distribution, and any difference you see may be the difference between those two things rather than between before and after. Recommendation

So the baseline runs the frozen panel, at the same run counts, over a declared window, before the correction is made. That ordering is the whole constraint, and it collides with the instinct to fix a wrong fact the moment you find it. Where the error is severe enough that waiting is not defensible - anything carrying an escalation flag, most obviously - correct it immediately and accept that you have given up the measurement. That is a legitimate trade and worth naming out loud, because the alternative is a team that quietly delays a compliance fix to protect a study.

Declare the baseline window in advance and record it. A baseline that ran until the numbers looked stable is a baseline chosen by its results.

Repetition, and Where the Run Counts Come From

Repetition is not thoroughness, it is the measurement. The same prompt on the same engine can return a different answer between runs, so a single post-correction run tells you what happened once. The variation page covers what has been measured about that instability, at the scope of the studies that measured it.

The published run-count guidance this cluster inherits comes from one preprint, and it is worth knowing exactly what it counted. It comes from Julius Schulte, Malte Bleeker and Philipp Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv 2604.07585, submitted April 8, 2026 - a preprint, not peer-reviewed), which ran across four engines and four commercial verticals in a German-language Swiss market and proposed at least seven runs per prompt-engine cell per day where brand detection is the outcome, and at least eight where source-level coverage matters. Our shortlist measurement page sets out the operational interpretation and the practitioner caveat: the dataset that adopted those numbers states plainly that its own cohort does not independently prove seven and eight are universally sufficient.

Here is the part that matters for this page, and we would rather state it than let it sit unnoticed. That guidance was derived on brand detection and source coverage. This page scores something different: whether one specific factual claim is stated correctly. We located no study establishing how many runs per prompt-engine cell are needed to characterise claim-level factual correctness, and there is no reason to assume the two outcomes need the same number. A binary fact may well be more stable than a brand's presence in a list, or less - we located no measurement either way.

So the practical position is this, and it is ours rather than anyone's finding. Recommendation Treat the published counts as a floor carried across from a different outcome. Run at least that many per prompt-engine cell per day in both the baseline and the observation window, keep the counts identical across the two, and state in your write-up that the figure was transferred from brand-detection guidance rather than derived for factual scoring.

If you think the baseline may be too unstable for that floor, predeclare an instability rule before collection rather than deciding afterwards. Recommendation Set the trigger in advance - a spread across baseline runs beyond a stated width, for instance. If the baseline trips it, increase the daily count and restart or extend the baseline so that the final baseline and observation windows use the same count, then report the rule, the trigger and which runs you retained. Do not raise the count merely because a first baseline estimate looks inconvenient: a sample size chosen after seeing the numbers is a design decision made by the results, and it leaves the two windows unequal.

The Controls You Can Have, and the One You Cannot

You can build a reference series. You cannot build a control group, and the difference is not pedantic.

Control prompts and unchanged facts. Alongside the corrected claim, run prompts touching facts you are deliberately not changing - ideally on the same product, of the same kind, and correct throughout. Recommendation They give you a picture of how much the panel moves on its own over the same window.

Decide in advance how that series changes what you write, because otherwise it becomes a test you never specified. Recommendation Predeclare the contrast as a descriptive one: report the baseline-to-observation percentage-point change for the corrected claim beside the corresponding change for each reference claim. Do not subtract them into a single estimate - that arithmetic looks like an effect size and is not one. If the corrected claim moved in a similar direction and by a similar magnitude to the reference series, describe the result as having occurred amid comparable background movement. If it did not, describe it as distinct from the observed reference movement. Both descriptions keep the non-causal limitation attached.

Note what that phrasing withholds. The reference series can weaken or strengthen an interpretation; it cannot classify a change as noise, because this design specifies no contrast, no uncertainty treatment and no threshold that would support that word. Comparable movement in unchanged claims weakens any reading that treats the corrected claim's movement as correction-specific. That is as far as it goes.

What the series is not. It is not a control group. Nothing is randomised, no unit was assigned to a condition, and there is no counterfactual - you cannot observe what the corrected claim would have done had you left it alone. A reference series bounds your interpretation. It does not license a causal one.

One change, or several. A one-change design corrects exactly one source, at a recorded timestamp, with everything else held as constant as you can manage. It is slower, and what it buys is temporal separation and a clean event record - not attribution. With no counterfactual available, a later change still cannot be attributed to that edit: recrawl timing, answer variability, model changes and other sources changing in the same window are all uncontrolled. A one-change design supports a temporal association. A multi-source design corrects everything at once, which is the faster route to a fixed estate and removes even that separation, leaving no basis on which any single edit could be told apart from the others.

Neither is wrong. The choice is a trade between speed and interpretability, and it belongs to whoever owns the commercial risk, not to the person running the panel. What is not acceptable is running a multi-source correction and reporting it as though one edit had been isolated. Decide before you start, and write the decision into the protocol.

The Observation Window, and What Happens After It

Fix the observation window in advance, measured from the verified source-correction date, and publish it with every figure you report. Recommendation A window that ends when the result looks good is not a window; it is a decision dressed as a measurement.

How long is a genuine question with no documented answer. No provider we located documents a cadence at which its index or its answers reflect changed page content. The timing statements providers do publish describe other things - requested recrawls, permission-signal propagation, the removal of a site's content from a feature - and none of them answers this question. What each of those four provider statements actually says, and what it does not cover, is set out on its own page.

Two consequences follow for the design. Recommendation

Sample repeatedly across the window rather than checking at the end. A single measurement at day 30 cannot tell the difference between a claim that corrected on day 3 and one that corrected on day 29, and those are different findings.

Keep measuring past the first corrected observation. A corrected answer is not a terminal state. A claim that reads correctly in one run can revert in the next, and a series that stops at the first good result will report a correction that may not have held. Regression is a finding, and a protocol that cannot observe it is not measuring persistence - it is measuring the moment you decided to stop.

Scoring at Claim Level

Score the disputed claim, not the answer. Everything else in the response - tone, completeness, whether a competitor appears, whether the answer is any good - is out of scope for this measurement. Recommendation

Each run gets one of three claim values, plus a run status where the run itself did not produce a scoreable answer.

Scoring values, per run
ValueMeaning
CorrectThe claim is stated, and it matches the authoritative value recorded in the incident
IncorrectThe claim is stated, and it does not match
Not addressedThe answer does not state the claim either way
No generated answerRun status, not a claim value - the surface returned links or declined to generate
Run voidRun status - an error, a timeout, a session that failed its own condition schema

"Not addressed" is not "correct," and treating it as correct is the most consequential scoring error available here. An answer that stops mentioning your pricing entirely has not been corrected; it has stopped answering. Folding those runs into the correct column produces a rising line that means nothing, and it is an easy fold to make because a wrong claim did disappear.

Four rules keep the scoring defensible. Recommendation

  • Score against the recorded authoritative value, not against what the scorer remembers the price or tier to be.
  • Score from exported text, not from the live interface, so the same artefact can be rescored later and by someone else.
  • Double-score a sample and record the disagreement rate. If two people disagree on 15% of runs, your headline figure carries that uncertainty and the write-up should say so.
  • Record reclassification. An incident logged as unsupported or unverified that gets adjudicated mid-study becomes a different type, and the date it changed matters to the series.

Reporting the Result

The reportable sentence has a fixed shape, and the shape is the safeguard: on engine E, over window W, the disputed claim was stated correctly in X of Y scoreable runs, against Z of Y in the baseline window. Recommendation

Everything that makes it defensible is in that sentence - the engine, the window, the numerator, the denominator, and the comparison. Everything that would make it indefensible is absent from it: there is no verb claiming the correction did anything.

Around that core, five reporting rules. Recommendation

  • Report per engine. Never pool, never average across engines, never present a single headline number covering several.
  • Reconcile the denominators. Void runs and no-answer runs are reported, not silently dropped, so that scoreable runs plus unscoreable runs equal runs attempted.
  • Report the request stages separately, on their own denominator of requests submitted.
  • Report the reference series alongside the corrected claim, so a reader can see how much the panel moved on its own.
  • Use no causal verb. Not "the correction fixed the answer," not "updating the pricing page resolved it." The claim was stated correctly in X of Y runs after the correction. That is the finding.

A result where nothing changed is a result, and it is reported in exactly the same form. There is a real temptation to leave a null out of the deck, or to re-cut it until something moves. A series that did not shift can be the most useful thing you learn in a quarter, because it redirects effort away from a source that is not doing what everyone assumed.

And where you cannot account for what you observed, write that. Name what you saw, what you could not observe, and what evidence would have distinguished the possibilities. "The claim was stated correctly in 22 of 30 runs from day 9, we observed no crawler fetch or displayed citation for the corrected page in that period, and we cannot account for the change" is a complete and honest finding. Inventing a mechanism to close the paragraph is not.

What This Design Can Never Establish

Four things sit permanently outside what a vendor-run correction panel can show, and a write-up that does not name them will be read as claiming them.

It cannot produce a prevalence, an error-type distribution, or an engine-accuracy ranking. The panel was seeded from answers already known to be wrong, which means it is a purposively selected cohort. Its ratios look exactly like rates and are not: a set assembled because it contained errors cannot tell you how often errors occur, which kinds are most common, or which engine gets things right more of the time. If someone asks which engine is most accurate about your product, this instrument cannot answer, and no amount of running it for longer will change that.

It cannot establish cause. There is no randomisation, no assignment and no counterfactual, so the ceiling on inference from an uncontrolled before-and-after is a temporal association: the claim was stated correctly more often after the correction than before, over these windows, on this engine. A before-and-after screenshot pair is weaker still - two single observations, one from each side of a change, with no distribution behind either.

A corrected answer does not confirm that your correction worked. The claim may have changed for reasons you did not observe and could not see: a third-party page updated in the same period, an index refresh, a model change, or ordinary variation. This is why the reference series exists and why the request stages, crawler-fetch observations and displayed-citation observations are recorded separately - not to prove the link, but to show a reader how much of it is observed and how much is assumed.

And the platform reporting available does not close the gap. Google's generative AI performance report in Search Console includes impressions for the following generative AI capabilities on Google Search: AI Overviews / AI Mode, broken down by page, country, date and device, with All dates are in Pacific Time Zone (PT) and the usual data limitations (1,000 row limitation, time period, etc); it also states that Search Console doesn't include data from experiments in Search Labs, as these experiments are still in active development. Documented The report's own listed dimensions - impressions by page, country, date and device - do not include prompt text or answer text; Google's help page does not itself state that this content is withheld, but nor does it describe surfacing it anywhere. It can tell you that a page drew impressions in these features. It cannot tell you whether an answer was corrected, which means the measurement stays manual.

One smaller limit worth writing into the method: where model and version are not exposed on a logged-out consumer interface, the observation date is the scope anchor for everything you report. A result belongs to the dates it was collected on, and to nothing wider.

Want a correction panel that produces a defensible number rather than an argument?

The design work is mostly decisions - which claims, which engines, how long, and what you are willing to leave uncorrected while you measure. Book a Session.

Sources

Sources: Provider documentation, quoted as read on September 1, 2026. Google Search Console Help, Generative AI performance report (Search) - the dimensions, the Pacific Time statement, the 1,000-row limitation and the Search Labs exclusion quoted above. It exposes no prompt or response text. The page states that as of August 31, 2026, we've rolled out these insights to all websites worldwide, with no contradictory rollout language found elsewhere on the page on re-check; it is cited to the review date because help documentation changes without notice. Run-count guidance is not ours and is not provider documentation. It originates in Julius Schulte, Malte Bleeker and Philipp Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)" - arXiv 2604.07585, submitted April 8, 2026, a preprint carrying no journal or proceedings reference and therefore not peer-reviewed - run across four engines and four commercial verticals in a German-language Swiss market, proposing at least seven runs per prompt-engine cell per day for brand detection and at least eight where source-level coverage matters. It was adopted by a practitioner dataset whose publisher states that its cohort does not independently prove those counts are universally sufficient; that operational interpretation and caveat are set out on our shortlist measurement page. The transfer to claim-level factual scoring is ours, and is an extrapolation: that guidance was measured on brand detection and source coverage, and we located no study establishing the run count required to characterise whether a specific factual claim is stated correctly. On the central question of this page: we ran a deliberate search across two research passes - formulations covering controlled and randomised designs, field experiments, before-and-after content interventions and documented vendor case studies - and located no published peer-reviewed paper and no preprint that intervenes on a web source and measures whether an AI answer subsequently changed. That search was re-run on the review date with the same result. What exists instead is practitioner commentary describing informal experiments without disclosed designs. This is a statement about what a literature search returned, not about what has been done: a controlled version of this test may exist privately or remain unpublished, and none was available to support this protocol on the review date. Accordingly, no design on this page is presented as establishing cause, and none of the protocol, run counts, scoring values, window rules or reporting format is documented by any provider or platform - they are CoreAEX's, and they are labelled as such throughout.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.