Two labels mark evidence boundaries on this page. Documented marks a platform definition or behaviour directly supported by the linked provider documentation. Recommendation means a workflow, threshold or decision rule that CoreAEX prescribes and no source specifies. Published research is attributed in the prose with its venue, design and scope, and is not tagged - it is evidence about a method, not documentation of how a product behaves. Untagged text is ordinary explanation or a conclusion following from something already labelled.

You have a pattern that survived. It came out of a backward trace, it held up against a comparison group, and it repeated on a cohort that played no part in finding it. There is a campaign budget and a quarter to spend it in, and the obvious move is to build the campaign around the pattern and watch the dashboard.

A dashboard cannot distinguish the campaign working from the campaign doing nothing, because it has nothing to compare against. Whatever the number does next, something else was also happening - seasonality, a competitor's pricing change, the two enterprise deals that were always going to close in March. The page on patterns and artefacts ends by saying that what converts an association into something you can act on is a prediction, written down before the period it applies to, with a comparison built in. This page is how you build that.

You now have two different questions, and the rest of this page depends on keeping them apart. First: does the frozen association recur in an untouched cohort? Second: if you build a campaign from it, does assigning that intervention change the defined outcome? The first is prospective replication; the second is causal evaluation. A later cohort can answer the first without answering the second.

And the campaign is not a separate step from the test. The default sequence is to run the campaign and then work out how to measure it. Designing the measurement before the spend often costs little relative to the campaign itself - though it is not free, since it can require analyst time, instrumentation, a negotiated holdout and sometimes a commercial concession - and it prevents an expensive ambiguity later, because the decisions that make a result readable have to be made before the spend starts rather than reconstructed from it.

Two tests, and they are not interchangeable

Pre-declaration improves both of the questions above. It does not make them the same question, and collapsing them is the most common way a good pattern turns into an overclaim.

Prospective replication and intervention effect
Test A - prospective replicationTest B - intervention effect
What you doFreeze the original exposure, outcome, cohort definition, exclusions and contrast, then apply them unchanged to an untouched mature cohortAssign a campaign, or construct an explicit counterfactual, and estimate the difference in a pre-declared outcome
What it asksIs the observational association stable enough to recur?Did acting on the pattern change the outcome, under this design's assumptions?
What it does not doEstablish why it recurs. A recurring association is still an associationTell you the association was real beforehand - you can get an intervention effect from a pattern that never replicated
Who owns it hereThe page on patterns and artefacts, which sets out the holdout and later-cohort rulesThis page, from here on

Everything below is Test B. It assumes you have a pattern worth spending money on, and it is about designing the spend so that the quarter can tell you something.

Why the writing-down has to happen first

The distinction between a finding and a hypothesis is not about how convincing the pattern looks. It is about when the hypothesis was fixed relative to the data.

Clinical research states this more plainly than marketing does, and it is worth borrowing the language. The FDA's guidance on statistical principles for clinical trials, adopted from ICH E9, distinguishes confirmatory trials from exploratory ones. On exploratory work it says:

"Tests of hypothesis may be carried out, but the choice of hypothesis may be data dependent. Such trials cannot be the basis of the formal proof of efficacy, although they may contribute to the total body of relevant evidence."

"The choice of hypothesis may be data dependent" describes a backward trace exactly. You looked at the data, and the data suggested what to look at. That is a legitimate way to generate a hypothesis and a poor way to confirm one - and note the second half of the sentence, which is the part that gets dropped: exploratory work still contributes to the body of evidence. It is not worthless. It is a different category.

The guidance's rule for the other category is short:

"Only results from analyses envisaged in the protocol (including amendments) can be regarded as confirmatory."

Two things about how this is being used. It is regulatory documentation governing clinical trials submitted for approval, quoted here as a definitional standard and an analogy - no regulator has ruled on marketing analytics, and nothing on this page should be read as though one had. And the parenthesis matters: including amendments. Pre-declaration is not a prohibition on changing your mind. It is a requirement that changes be recorded when they are made rather than discovered afterwards in the results.

The same idea has a name in research practice. Nosek, Ebersole, DeHaven and Mellor define it in PNAS (2018) as "define the research questions and analysis plan before observing the research outcomes - a process called preregistration," which "distinguishes analyses and outcomes that result from predictions from those that result from postdictions." Predictions and postdictions is the whole distinction in two words. This page cites that paper for the definition and for nothing else; whether the practice measurably improves reproducibility is a separate empirical question that we make no claim about here.

There is a published objection to how that argument is usually made, and it sharpens the practical advice rather than weakening it. Alison Ledgerwood, replying to Nosek and colleagues in the same journal (PNAS, 2018), argues that discussions of preregistration "seem to conflate the goal of theory falsification with the goal of constraining type I error," and that this "masks a crucial distinction between two types of preregistration: preregistering a theoretical, a priori, directional prediction (which serves to clarify how a hypothesis is constructed) and preregistering an analysis plan (which serves to clarify how evidence is produced)." Her summary is two sentences long: "Preregistering theoretical predictions enables theory falsifiability. Preregistering analysis plans enables type I error control."

Note that this is a different cut from the one above, and the two should not be run together. Test A and Test B are two empirical questions. Ledgerwood's distinction is between two purposes of writing something down in advance, and it applies inside either question. Writing down a confident prediction does not, by itself, do anything about the false-positive rate; specifying the analysis is what does that.

Two consequences land directly on the form in the next section. The first is the failure mode she names - "a researcher attempting to control type I error records careful predictions but omits or only loosely specifies a preanalysis plan" - which describes most marketing pre-declarations exactly: a bold hypothesis sentence followed by vague analysis detail. The second is more encouraging. She notes that conflating the two "implies, erroneously, that preanalysis plans can only help control type I error when research is in a prediction-making/theory-testing phase," when "in fact, preanalysis plans can also be useful in the question-asking/discovery/theory-building phase." You do not need a confident prediction to benefit from pre-specifying the analysis. A team that only has a hunch still gets the type I error benefit from settling the outcome, the population and the analysis in advance. Her letter is an argument about research practice in psychology, and reading it across to a marketing test is an application of the distinction rather than a finding about marketing.

Why this bites harder on a backward trace than on most analyses. The trace's candidates were produced by searching - across dimensions, cross-tabs, lookback rules and exclusions. The multiplicity section of the artefact page sets out what that search does to an ordinary significance claim. Writing the prediction down before the next period is what converts the surviving candidate from something the search produced into something the next quarter can actually test.

The pre-declaration, as a form to fill in

Everything above is procedural, so here it is as a form rather than an argument. Recommendation Fill in all six fields before the period starts, save it somewhere with a date on it, and share it with whoever will read the result. Fields left blank are the ones that will be filled in afterwards by whatever happened.

The first row is a prediction; the other five are an analysis plan. They do different jobs, in Ledgerwood's sense, and the second job is the one that protects the false-positive rate. A form with a confident first row and four vague ones is the common failure, and it is the one worth checking for before the period starts rather than after.

Test pre-declaration - complete before the period begins
FieldWhat goes in itWhy it has to be now
The predictionOne sentence, directional and specific: "accounts in segment S exposed to X will reach the defined SQL stage at a higher rate than comparable accounts that are not"A vague prediction is compatible with any result, which is the same as no prediction
The populationWho is eligible to be in this at all - segment, market, product line, size band - and who is excluded, with the exclusion rule written outExclusions chosen after seeing the data are the cheapest way to manufacture an effect
The comparison groupThe specific units that will not receive the intervention, and how they were assignedWithout a concurrent control or a credible counterfactual, the result is an uncontrolled before-and-after. It can describe change over time; it cannot isolate the campaign's effect
The outcomeOne primary measure, with its definition, its window and its denominator. Secondary measures listed and labelled secondaryChoosing the outcome afterwards from several measured means picking the one that moved
The difference that mattersThe smallest effect you would act on, in the outcome's own units - and the arithmetic showing whether your volume can detect itWithout it you cannot tell an inconclusive result from a null one
The stopping ruleWhen the test ends: a fixed calendar date or volume, or prospectively planned interim looks with stopping boundaries and an analysis that controls the resulting error rateStopping because an unplanned interim result looks interesting is a decision made in view of the answer. Planning the interim looks in advance is not the same thing

The last row generalises, and it is the discipline the whole page rests on. Anything chosen after seeing which answer it gives - the stopping point, the exclusion, the outcome measure, even which confidence interval to report - is a decision made in view of the result. Each one individually feels like judgement. Together they are how an underpowered test produces a publishable-looking number.

What you can actually compare against, in B2B

The clean answer is randomisation at the level of the buyer. Three things stand in the way of it in B2B, and none of them is a methodological objection you can argue past. Marketing reaches accounts through channels that do not respect a buyer-level split. Sales teams talk to whoever raises a hand. And withholding a legitimate approach from a named target account is a commercial decision rather than a measurement one. Saying so plainly is more useful than pretending otherwise.

What remains is a short list, in descending order of what it can establish.

  • Randomised assignment at a level you can control - geography, territory, segment, industry vertical, or a list of target accounts split at random before the campaign is built. When the assignment is implemented as planned, outcomes are analysed according to the assignment rather than to who actually received the treatment, clustering is handled, and material interference or differential drop-out is either absent or addressed, this supports a causal estimate for the units randomised and the intervention studied. Each of those conditions is a place the design can quietly stop being an experiment.
  • A holdout that was not randomised - a region or segment left out for operational reasons. Weaker, because whatever made it the one left out may also relate to the outcome, but honest if the reason is stated.
  • A modelled counterfactual - constructing what the treated units would have done, from control units or from their own history. Validity rests on assumptions the method cannot check for you. See below.
  • A later cohort, unrandomised - the weakest of the four, because anything else that changed between periods is inside the comparison.

Recommendation Take the strongest of these that your commercial constraints actually permit, and record which one you took in the pre-declaration - because the strength of the claim you can make at the end is fixed by this choice, not by how the result turns out. A team that runs the fourth option and reports it as the first has not run a test; it has run a rollout with a paragraph attached.

What a market-level test establishes

One practical randomised design, where exposure can be controlled geographically, assigns markets rather than individuals. Jon Vaver and Jim Koehler set out the approach for Google in 2011 - company research rather than a peer-reviewed paper, with no journal, no volume and no DOI, which is worth knowing before leaning on it. Their description of the design is the useful part:

"In these experiments, non-overlapping geographic regions are randomly assigned to a control or treatment condition, and each region realizes its assigned condition through the use of geo-targeted advertising."

They are direct about why the randomisation is what matters: "Randomization is an important component of a successful experiment as it guards against potential hidden biases," because "there could be fundamental, yet unknown, differences between the geos and how they respond to the treatment."

Here is the consequence that matters most for B2B, and it is a property of the design rather than something the geo paper states. If geographies are randomised, they are the assignment clusters, and the analysis has to account for the correlation between observations inside each one. A dozen geographies is a very small cluster count, and that sharply constrains degrees of freedom, power and control of the false-positive rate.

What does not follow is that the accounts and deals inside those geographies contribute nothing. They do. Precision in a clustered design depends on the number of clusters, their sizes, the variation within and between them, the correlation inside a cluster, how balanced the assignment is, and which analysis you run. Saying "twelve markets is twelve units, whatever is inside them" is as wrong in one direction as counting every deal as an independent observation is in the other.

The methods literature is blunt about what a small cluster count does. Leyrat, Morgan, Leurent and Kahan ran a simulation study of twelve analysis approaches for cluster-randomised trials with forty or fewer clusters and continuous outcomes (International Journal of Epidemiology, 2018). They report that "mixed models and GEEs can lead to inflated type I error rates with a small number of clusters," that corrected approaches "had low power (below 50% in some scenarios) when fewer than 20 clusters were randomized, with none reaching the expected 80% power," and - the sentence a marketing team should sit with - that "for a given number of clusters, increasing the cluster size may not always be sufficient to reach an 80% power, and thus the randomization of a larger number of clusters is preferable whenever feasible." Their scope is simulated continuous outcomes in a clinical-trials context rather than pipeline data, and what transfers is the structural point rather than any particular number.

Recommendation So plan power for the clustered design as a whole: account for the number and sizes of the clusters, the intracluster correlation, the balance of the allocation and the analysis method, rather than treating every account or deal as independent. Do not reuse the individual-level two-proportion arithmetic from the starting-point page unless individuals are what you randomised - and if the cluster count is small, choose the analysis method deliberately, because with few clusters the uncorrected default is the one that inflates the error rate.

The paper names its own practical limits too, and they transfer directly:

"Ad serving inconsistency is a concern due to finite ad serving accuracy, as well as the possibility that consumers will travel across geo boundaries."

In B2B that second concern is sharper than in consumer advertising. A buying committee spans locations, a company headquartered in a control market has staff in a treatment market, and the account is the buying unit while the assignment was geographic. Recommendation Where accounts rather than places are the thing you are trying to influence, randomise the accounts - split a target list at random before the campaign is designed, and treat the account as the unit throughout. It is the same design with a unit that matches the buying reality, and it reduces the mismatch between the assignment unit and the buying account. It does not eliminate spillover: shared buyers, agencies, partners, public content, sales teams and word of mouth all cross account lines.

None of this is available to every company. A single-market business with one sales team has no geography to split and may have too few target accounts to split either, and the honest answer there is in the last section rather than in a smaller version of this one.

Modelled counterfactuals, and what they assume

When nothing can be randomised, the remaining option is to model what would have happened. These methods are legitimate and well documented, and each one rests on an assumption it cannot verify from your data - which is not a criticism of the methods but a description of what they are.

Bayesian structural time series, the approach behind Google's CausalImpact, builds a counterfactual for a treated series from control series that were not themselves treated. Brodersen, Gallusser, Koehler, Remy and Scott state the condition in their peer-reviewed paper (Annals of Applied Statistics, 2015): "This approach assumes that covariates are unaffected by the effects of treatment." The package documentation repeats it with more force: "all of the above inferences depend critically on the assumption that the covariates were not themselves affected by the intervention." Documented It also states a second condition that matters over a long campaign, in its own words: "As long as the control series received no intervention themselves, it is often reasonable to assume the relationship between the treatment and the control series that existed prior to the intervention to continue afterwards."

Read those against a real marketing programme. If your campaign runs in one segment while brand activity, a pricing change or a partner push touches the segments you are using as controls, the assumption is under threat - though not automatically broken. What matters is whether the other activity affected the control series, or changed their relationship to the treated series in a way the model does not capture. If it did, the method will still return a number with an interval around it.

But it is not true that nothing can warn you, and the documentation is the reason to say so. Its FAQ names four checks: reason explicitly about whether each covariate could have been affected by the intervention; "plot all covariates and do a visual sanity check"; run the analysis "on an imaginary intervention" in the pre-period and confirm it finds no significant effect, since "counterfactual estimates and actual data should agree reasonably closely"; and "when presenting or writing up results, be sure to list the above assumptions explicitly, including the priors." Documented

Note what those checks do and do not reach. The placebo run tests pre-period predictive fit. The covariate question - the identifying assumption everything rests on - is addressed by reasoning and a visual inspection rather than by a test the software performs. The documentation is candid about this, opening the answer by calling assumption verification "the elephant in the room with any causal analysis on observational data" and offering the list as "a few ways of getting started." Diagnostics can expose some failures - contaminated controls, poor pre-period prediction, a placebo period that shows an effect. They cannot establish that the assumption holds.

Synthetic control builds the counterfactual as a weighted combination of untreated units. Abadie, Diamond and Hainmueller's peer-reviewed paper (Journal of the American Statistical Association, 2010) sets the condition on when it should be used at all: "In some instances, the fit may be poor and then we would not recommend using a synthetic control." They also state two assumptions worth carrying - "the usual assumption of no interference between units," and that "the intervention has no effect on the outcome before the implementation period." A leaked launch, a pre-announcement or a sales team that started selling the story early breaches the second one only if it moved the outcome before the stated implementation date - which is worth checking rather than assuming in either direction.

Marketing mix models make the same bet

The same logic applies one level up, to the models that attribute across all channels at once. Google's own documentation for Meridian is direct about it, in the page setting out the method's required assumptions: "In practice, it is difficult to know whether all of the confounding variables are measured because it is purely an assumption, and there is no statistical test to determine this from your data." Documented That is a vendor describing the limits of its own tool, dated 8 July 2026 and reviewed here on 4 September 2026.

Recommendation Stating the identifying assumption is necessary and not sufficient. Write it into the pre-declaration in your own words; explain why it is plausible for your business; examine the pre-period fit and the placebo behaviour; test how much the answer moves when you change which controls are included; and document anything you know about the period that could have contaminated them. If the assumption still cannot be supported after that, report the statistical estimate as a statistical estimate and do not present it as a defensible causal effect.

Read the design for power before you run it

This is the step that gets skipped, and skipping it makes a small test easy to overinterpret. An imprecise experiment can still be informative when the estimate and its uncertainty are reported honestly. What it cannot survive is being read as though it were precise.

Andrew Gelman and John Carlin set out the consequence in a peer-reviewed methods paper (Perspectives on Psychological Science, 2014). They define two errors that ordinary power calculations ignore - a Type S error, "the probability that the replicated estimate has the incorrect sign, if it is statistically significantly different from zero," and a Type M error, "the factor by which the magnitude of an effect might be overestimated." Their conclusion is the sentence to keep:

"In fact, statistically significant results in a noisy setting are highly likely to be in the wrong direction and invariably overestimate the absolute values of any actual effect sizes, often by a substantial factor."

Turn that around and it stops being abstract. A small B2B test that returns a large, statistically significant effect has not necessarily found a large effect. In a noisy design, a large estimate is the only kind of estimate that could have cleared the threshold - so the size of the number is partly a property of the filter it passed through. The paper's examples are drawn from psychology and medicine, and the statistical argument is general; the specific exaggeration factors they compute belong to their examples and do not transfer.

Recommendation So do the arithmetic before the period, on the unit you are actually randomising, and record the answer in the pre-declaration next to the difference that matters. The sufficiency section of the starting-point page has the calculation. Two things travel with it. A required sample size is specific to the assumptions you fed it - the baseline rate, the allocation, the significance level, the power, the test - and reaching that number does not guarantee a distinguishable result. And if the arithmetic says the design has too little power to reliably distinguish a difference you would act on under the stated assumptions, you have learned that before spending the budget rather than after, which is the entire reason for doing it first.

When no readable test exists

Sometimes there is nothing to randomise, no usable control series, and not enough volume to detect anything you would act on. That is a result about your measurement situation, and it has three honest responses that are all better than running an unreadable test and reporting its number.

Run the campaign, and say what it is. A campaign built on a surviving pattern is a reasonable commercial bet. Report it as a bet whose effect you will not be able to isolate, and record the pre-declaration anyway - because a prediction written down before an unreadable period is still a prediction, and a year of them starts to be informative in a way that a year of post-hoc explanations never becomes.

Change what you can measure rather than what you can conclude. If the account list is too small to split this quarter, it may not be next year; if no control segment exists because the campaign runs everywhere, a staggered launch may create comparative variation. Staggering is not a free control, though: how it reads depends on timing, spillover and the analysis you plan for it, it can introduce calendar confounding, anticipation and carryover, and it can carry a real commercial cost. Assess whether a randomised or otherwise justified stagger is operationally acceptable rather than assuming one is available. The most useful output of a failed power calculation is a design change, not a caveat.

And make the tests cheaper rather than the claims looser. The published research on advertising measurement has largely moved in this direction. Johnson, Lewis and Nubbemeyer's peer-reviewed work on "ghost ads" (Journal of Marketing Research, 2017) is explicitly about reducing the cost of experimentation rather than improving the models that substitute for it: their method, they report, lets "advertisers can measure ad lift just as precisely while spending at least an order of magnitude less." Their demonstration is one online retailer's display retargeting campaign, which is neither B2B nor a pipeline outcome, and the specific lift figures they report belong to that campaign; what carries across is the direction of the answer, not the numbers. It is also not a general B2B recipe: the method needs ad-platform auction and counterfactual-exposure capabilities that many teams cannot reproduce. When experiments look unaffordable, the productive question is how to make one cheaper, not which model can stand in for it.

What none of these is: a licence to report the campaign's period-over-period change as its effect. Measuring content-market fit owns the bar that applies here, and the bar does not move because the test was hard to build.

What each design licenses

The design fixes the sentence, and it fixes it before the result exists. That is the point of settling it in advance.

Test design and the strongest claim it supports
What you ranWhat you can sayWhat you cannot
Randomised assignment of markets, segments or accounts - pre-declared, implemented as planned, analysed by assignment, clustering handled, and material interference or differential drop-out absent or addressed"Assigning X to these units changed the pre-declared outcome by this much, with this interval" - a causal statement at the level the assignment happenedThat the effect holds at a different level, in unassigned segments, or for units unlike the ones randomised - and the causal reading itself if any of the conditions in the first column failed, since each one is a place the design stops being an experiment. It also does not survive scaling automatically: what scaling changes is the next decision, not this one
Randomised assignment, not powered for the difference you would act on"This design was not planned with adequate power to detect a difference of D reliably under the stated assumptions. Here is the treatment-effect estimate and its interval, the planned and achieved sample structure, and a design analysis based on effect sizes justified from outside this result"Treat a non-significant result as evidence of no effect, or treat the size of a threshold-crossing estimate as self-validating. Low-power filtering can exaggerate significant estimates, so assess Type S and Type M risk against externally justified effect sizes
Non-randomised holdout, reason for the split stated"The treated and untreated groups differed by this much, and here is why the split happened and what else may differ with it"A causal claim. The thing that decided the split may also relate to the outcome
Modelled counterfactual with its assumption examined and stated"Under the assumption that the controls were unaffected and their relationship to the treated series was stable, the estimated impact is this"That the assumption has been proved. Diagnostics can expose some failures, but they cannot establish that the identifying assumption holds
Period-over-period, no comparison"This is what happened during the campaign"That the campaign did it

Every row above is a real thing to report, and the second row is the one worth defending internally. "We ran the test properly and it could not settle the question" is an uncomfortable slide and an honest one, and it is the row that tells you what to build differently next quarter. A pattern that has been through a trace, a comparison group, an out-of-sample check and a pre-declared test has been examined more carefully than almost anything else in a marketing report - and what happens to it as you scale is a separate question again.

Not sure whether your next quarter can carry a readable test? The two questions that settle it are what you are able to randomise and what difference would change a decision. Both are answerable before any budget is committed, and answering them first costs less than discovering the answer from an unreadable result.

Book a Session


Sources

Sources. Pre-specification: International Conference on Harmonisation, "E9 Statistical Principles for Clinical Trials," guidance for industry issued by the U.S. Food and Drug Administration (CDER and CBER), September 1998. The confirmatory sentence is in section V.A, "Prespecification of the Analysis" (5.1); the exploratory-trial sentences are in section 2.1.3. This is regulatory documentation governing clinical trials submitted for approval, quoted here as a definitional standard and an analogy. No regulator has ruled on marketing analytics, and nothing here should be read as though one had. Preregistration: Brian A. Nosek, Charles R. Ebersole, Alexander C. DeHaven and David T. Mellor, "The preregistration revolution," PNAS 115(11), 2018, pp. 2600-2606, DOI 10.1073/pnas.1708274114 - peer-reviewed, published in a Sackler Colloquium special feature. The two sentences quoted here are the paper's definitional statements, and the paper is cited for that definition alone. This page makes no claim about whether preregistration measurably improves reproducibility, which is a separate empirical question and not one this page needs. The published reply is cited alongside it: Alison Ledgerwood, "The preregistration revolution needs to distinguish between predictions and analyses," PNAS 115(45), 6 November 2018, pp. E10516-E10517, published online 19 October 2018 - a Letter rather than a research article, arguing that discussion of preregistration conflates theory falsification with type I error control. Its quotations here were taken from the published PDF. Its subject is research practice in psychology; applying its distinction to a marketing test is an application rather than a finding about marketing, and the page says so where it is used. Geo experiments: Jon Vaver and Jim Koehler, "Measuring Ad Effectiveness Using Geo Experiments," Google Inc., 2011 - company research, not peer-reviewed, with no journal, volume, pages or DOI. Quoted for its description of random assignment of geographic regions and for its stated concerns about ad-serving accuracy and cross-boundary movement. Those concerns are not transferable to B2B without qualification - the page explains how the same risk mechanisms operate differently when the buying unit is an account spanning several locations rather than a consumer in one. The paper makes no statement about required sample size, number of geos or statistical power; the discussion on this page of how cluster count, cluster size, within-cluster correlation, allocation balance and analysis method jointly affect precision is CoreAEX's interpretation of the design and the methods literature, not a finding of theirs. Modelled counterfactuals: Kay H. Brodersen, Fabian Gallusser, Jim Koehler, Nicolas Remy and Steven L. Scott, "Inferring causal impact using Bayesian structural time-series models," Annals of Applied Statistics 9(1), 2015, pp. 247-274 - peer-reviewed; the covariate assumption is quoted from the paper and, in its stronger phrasing, from Google's CausalImpact package documentation, which is project documentation for a specific tool rather than peer-reviewed research. The four diagnostic checks quoted on this page are that documentation's own FAQ answer to the question "How can I check whether the model assumptions are fulfilled?", including its framing of assumption verification as "the elephant in the room." Alberto Abadie, Alexis Diamond and Jens Hainmueller, "Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California's Tobacco Control Program," JASA 105(490), 2010, pp. 493-505 - peer-reviewed. Note that the line about not recommending the method when pretreatment fit is poor or pretreatment periods are few belongs to the authors' later paper (American Journal of Political Science 59(2), 2015), not to this one, and is deliberately not used here. Google, Meridian documentation, "Required assumptions" - vendor documentation for a specific tool, page dated 8 July 2026 and reviewed 4 September 2026; a dynamic page that may change. Power, sign and magnitude: Andrew Gelman and John Carlin, "Beyond Power Calculations: Assessing Type S (Sign) and Type M (Magnitude) Errors," Perspectives on Psychological Science 9(6), 2014, pp. 641-651 - peer-reviewed. Their examples are from psychology and medicine; the statistical argument transfers, and the specific exaggeration factors they compute do not. Cheaper experiments: Garrett A. Johnson, Randall A. Lewis and Elmar I. Nubbemeyer, "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness," Journal of Marketing Research 54(6), 2017, pp. 867-884 - peer-reviewed; quoted from the published abstract. Its single empirical demonstration is one online retailer's display retargeting campaign, which is consumer-facing rather than B2B and measures site visits and purchases rather than pipeline; the lift figures it reports are deliberately not quoted here, because they belong to that campaign and would not transfer. Cluster-randomised designs: Clémence Leyrat, Katy E. Morgan, Baptiste Leurent and Brennan C. Kahan, "Cluster randomized trials with a small number of clusters: which analyses should be used?", International Journal of Epidemiology 47(1), 2018, pp. 321-331 - peer-reviewed. It is a simulation study of twelve analysis approaches across ninety scenarios, restricted to continuous outcomes and to trials with forty or fewer clusters, in a clinical-trials context; it is not a study of marketing experiments and reports no result about pipeline data. It is cited here for the structural point about small cluster counts, and the minimum-cluster figures it reports are described in the paper as suggestions from earlier work rather than as its own finding. Prospectively planned interim analyses: U.S. Food and Drug Administration (CDER and CBER), "Adaptive Designs for Clinical Trials of Drugs and Biologics: Guidance for Industry," final guidance, November 2019 - regulatory guidance, non-binding, and quoted here on the same analogical footing as ICH E9. It defines an adaptive design as "a clinical trial design that allows for prospectively planned modifications to one or more aspects of the design based on accumulating data from subjects in the trial," and defines "prospective" to mean "that the adaptation is planned and details specified before any comparative analyses of accumulating trial data are conducted." It is used on this page only to establish that a planned interim look is a different thing from an unplanned one, and it is equally clear that adaptive designs "should therefore address the possibility of Type I error probability inflation" and that a conventional end-of-trial estimate after an adaptive stop "would tend to overestimate the true population treatment effect." The selective-inference argument referred to throughout is set out and sourced on the page on patterns and artefacts, and the sample-size arithmetic on where to start a backward analysis; neither is re-derived here.

About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.