Two labels mark evidence boundaries on this page. Documented marks a platform definition or behaviour directly supported by the linked provider documentation. Any surrounding interpretation remains CoreAEX analysis unless separately attributed. Recommendation means a workflow, threshold or ownership rule that CoreAEX prescribes and no source specifies. Published research is attributed in the prose with its venue, design and scope, and is not tagged - it is evidence about a method, not documentation of how a product behaves. This page cites no platform documentation, so no Documented tag appears on it; that is expected rather than an omission. Untagged text is ordinary explanation or a conclusion following from something already labelled.
You ran the trace and something jumped out. Eight of your eleven wins touched the same page. Or they all came through the same channel, or every one of them had a security reviewer on the deal, or they clustered in a vertical nobody was targeting. It looks like a finding. You are about to put it in a deck.
Four distinct mechanisms can produce or exaggerate a striking pattern in a backward trace, and they compound rather than compete. Note "or exaggerate": a genuine association and an artefact can coexist, so the observed magnitude may be inflated even when the underlying association is not zero. Knowing which mechanism you might be looking at is the difference between a hypothesis worth testing and a budget reallocated on noise. Neither statistical significance nor a bigger sample settles it. Confidence increases when the pattern survives an appropriate comparison group and is then confirmed on data that was not used to discover it.
Four mechanisms, not one
These get lumped together as "bias" and treated as a single caution to be waved at in a methodology footnote. They are different problems with different fixes, and none is solved merely by adding more wins to the same biased, exploratory design. More observations improve precision and power only after the cohort, the comparison, the measurement and the hypothesis family have been defined.
| Mechanism | What creates or inflates the apparent pattern | What it responds to |
|---|---|---|
| Selection | Restricting the analysis to wins conditions on a shared downstream effect, which can associate causes that are unrelated across the full cohort | A comparison group from the eligible cohort |
| Survivorship | The records available to look at are the ones that survived a process that removed the rest | Recovering or accounting for what was removed |
| Multiplicity | A trace tests dozens of dimensions at once; some will look striking by chance alone | Counting the comparisons, then confirming the survivor out of sample |
| Halo | Knowing a deal was won colours how every other attribute of it is judged and recorded | Blinding, or measurement that predates the outcome |
Note the last column. Not one of these responds to "collect more wins." That is the section this page builds toward, because "we will know once we have a full year of data" is the reflex answer to all four, and it addresses none of them directly - volume becomes useful only once the cohort, the comparison and the hypothesis are already fixed.
Selection: the trace conditions on the outcome
The structural argument belongs to the pillar, which sets out why a comparison group drawn from the eligible, mature cohort is what makes a backward trace interpretable at all. The short version, because the rest of this page depends on it:
Hernán, Hernández-Díaz and Robins describe a causal structure shared by a family of biases - "conditioning on a common effect of 2 variables, one of which is either exposure or a cause of exposure and the other is either the outcome or a cause of the outcome" (Epidemiology, 2004, a theoretical paper using causal diagrams, with epidemiological examples). Closed-won status can have that shape whenever both marketing exposure and pre-existing account characteristics affect the probability of winning. Filter to the wins and you have conditioned on it, and associations can appear inside that filtered view that do not hold across the cohort.
What matters here is that selection is not a small distortion you can subtract afterwards. It changes what the numbers mean. A pattern produced this way is not a weak signal that more observations would sharpen: under a stable selection mechanism, adding more wins may make it more precise without making it more representative.
Survivorship: what a winners-only sample flatters
Survivorship overlaps with selection and is worth separating, because in this application it has a recognisable tendency - though not a universal one. Selection effects can move an estimate in either direction depending on the mechanism; what follows is about the direction this particular mechanism tends to produce.
Jerker Denrell's argument (Organization Science, 2003 - an analytical paper about organizational survival, not an empirical study of firms) is that "the organizations that can be observed at any point in time are the survivors of a selective process that has eliminated a large fraction of the underlying population," and that in consequence "in particular, risky practices, even if they are unrelated to performance in the full population of organizations, may seem to be positively related to performance in a sample of survivors." The direction is the useful part. He states it plainly: "observations of existing organizations will show that unreliable, uninformed practices and practices that involve concentrated resource allocation are superior to reliable, informed practices or practices that involve diversified resource allocation."
Read that against what a backward trace is usually used to justify. The conclusion a survivor-only sample is structurally disposed to produce - concentrate resources, the bold bet worked - is the same conclusion the meeting wants. That is not proof the conclusion is wrong. It is a reason to be suspicious of how easily it arrived. Applying an argument about organizational survival to a cohort of sales opportunities is an inference from a shared mechanism, not a documented finding about deals.
The mechanism is old and general enough to have been described in wartime. Abraham Wald's memoranda for the Statistical Research Group at Columbia, written during the war and reprinted by the Center for Naval Analyses in July 1980 as report CRC 432, set out the problem of estimating aircraft vulnerability when the only observable sample is the planes that came back. Wald states the information available - the total number of planes that took part, and, for each count of hits, the number of planes that received exactly that many hits and returned - and shows that it cannot be done: "it can easily be shown that it is impossible to estimate both [the hit-distribution] and [the survival probabilities] from the damage to returning planes only." Two caveats travel with that citation. The mathematical notation renders inconsistently across copies and is described here rather than quoted. And the popular version of this story - put the armour where the returning planes have no holes - is a later retelling rather than Wald's own wording. The memoranda are an estimation method, and Bill Casselman's account of them for the American Mathematical Society notes that "Wald says nothing about what the military should do to improve things," explaining that it was Statistical Research Group policy to answer the question asked rather than advise on applications.
The same structure has been demonstrated in a domain with far better data than any CRM. Brown, Goetzmann, Ibbotson and Ross, examining growth equity mutual funds (Review of Financial Studies, 1992), "analyze the relationship between volatility and returns in a sample that is truncated by survivorship and show that this relationship gives rise to the appearance of predictability," and present numerical examples showing the effect "can be strong enough to account for the strength of the evidence favoring return predictability." Note their exact wording: the appearance of predictability. Not weak predictability, and not a smaller effect than reported - an artefact of the truncation strong enough to account for the finding it produced.
And the mechanism is correctable, which is why this page is not an argument against looking backward. Denrell and Kovács (Administrative Science Quarterly, 2008, a simulation study) show that selective sampling of empirical settings distorts inference in two established research programmes - and close by saying they "discuss the implications of such selective sampling of empirical settings and suggest ways to correct for the bias." Correction is possible. It is just not automatic, and it does not happen by accident.
Multiplicity: how many things did you actually test?
This is the mechanism most often missing from marketing analysis entirely, and in a backward trace it is the largest of the four. The other three distort a comparison. This one manufactures comparisons.
Count what a trace actually examines. Channel, campaign, content asset, topic, entry page, buying role present on the deal, seniority of the first contact, number of contacts, industry, company size band, region, product line, deal size band, sales cycle length, quarter, rep, source of first touch, whether a particular page was viewed, whether a demo preceded a trial. Twenty dimensions is a modest trace. Cross two of them - channel by industry - and you have hundreds of cells. You did not run one test. You ran a great many, and then reported the one that looked like something.
Simmons, Nelson and Simonsohn put numbers on how fast this escalates (Psychological Science, 2011). Using computer simulations - 15,000 simulated samples, a two-condition design with 20 observations per cell, drawn from a normal distribution - they estimated how ordinary, undisclosed analytic choices change the probability of a false positive. Their term for those choices is worth quoting in full:
"The culprit is a construct we refer to as researcher degrees of freedom. In the course of collecting and analyzing data, researchers have many decisions to make: Should more data be collected? Should some observations be excluded? Which conditions should be combined and which ones compared? Which control variables should be considered? Should specific measures be combined or transformed or both?"
Their Table 1 reports the false-positive rate for each choice individually and in combination. At the p < .05 level - the table also reports p < .1 and p < .01 - "two dependent variables (r = .50)" produced 9.5%; "addition of 10 more observations per cell" 7.7%; "controlling for gender or interaction of gender with treatment" 11.7%; and "dropping (or not dropping) one of three conditions" 12.6%. Combining all four produced 60.7%.
Two things about that figure. The parenthesis in "(or not dropping)" is doing real work: the degree of freedom is the choice, exercised in either direction, not the act of dropping something. And these are simulation results under stated conditions, not an observed rate in any published literature or any marketing dataset. What transfers is the shape of the problem - four ordinary, individually defensible decisions taking a nominal one-in-twenty risk past one in two - not the number itself.
John Ioannidis frames the same problem as a property of a research environment rather than of an individual analysis (PLoS Medicine, 2005 - an essay presenting an analytical and simulation framework, not a measurement of any body of literature; a 2022 correction fixed a missing set of parentheses in one equation in Table 2 and changed nothing in the argument). Two of his corollaries describe a backward trace almost exactly:
"Corollary 3: The greater the number and the lesser the selection of tested relationships in a scientific field, the less likely the research findings are to be true."
"Corollary 4: The greater the flexibility in designs, definitions, outcomes, and analytical modes in a scientific field, the less likely the research findings are to be true."
A backward trace is high on both. It probes many relationships with little pre-selection, and it is unusually flexible in definitions - what counts as a touchpoint, which lookback window, how a contact maps to a buying role, which deals are "comparable." Ioannidis's own summary of what his framework produces preserves its status carefully, and so should any page citing it: "Simulations show that for most study designs and settings, it is more likely for a research claim to be false than true."
What to do about it, and the remedy that does not quite fit
The statistical literature has a standard answer to multiple testing. Benjamini and Hochberg proposed controlling the false discovery rate - defined as the expected value of the proportion of rejected null hypotheses that were erroneously rejected, with that proportion set to zero when nothing is rejected - rather than the probability of any false positive at all (Journal of the Royal Statistical Society: Series B (Methodological), 1995). Their summary describes it as calling "for controlling the expected proportion of falsely rejected hypotheses - the false discovery rate," and they prove a step-up procedure achieves it.
The scope condition on that proof is mandatory and is routinely dropped. Their Theorem 1 states: "For independent test statistics and for any configuration of false null hypotheses, the above procedure [the step-up procedure defined immediately above it] controls the FDR at q*." The 1995 result is for independent test statistics.
The dimensions in a backward trace are emphatically not independent. Channel correlates with content, because a channel delivers particular assets. Industry correlates with deal size, product line and sales-cycle length. Buying-role composition correlates with company size.
That does not mean multiplicity correction is unavailable under dependence, and it would be wrong to leave the impression that it is. Benjamini and Yekutieli extended the result (Annals of Statistics, 2001): they prove the original procedure "also controls the false discovery rate when the test statistics have positive regression dependency on each of the test statistics corresponding to the true null hypotheses," a condition their paper calls PRDS and which covers cases including comparisons of many treatments against a single control and multivariate normal statistics with a positive correlation matrix. And "for all other forms of dependency, a simple conservative modification of the procedure controls the false discovery rate" - a more conservative variant that holds under arbitrary dependence.
So the accurate position is narrower than "the correction does not apply here," and more demanding than "apply the correction." A backward trace does not automatically satisfy PRDS, and nobody has shown that it does. Using any of these procedures requires a defined family of tests and a stated dependence assumption, and a trace assembled by exploration has neither. Recommendation What is worth taking from this literature without a defined test family is the discipline underneath the formula:
- Write down how many dimensions you examined, including the ones that showed nothing. A finding reported without its denominator of comparisons is uninterpretable, and in our experience that denominator is rarely reported.
- Say which comparison you intended to make before you looked. One pre-specified comparison and one hypothesis generated afterwards are different objects, and only the first can be read at anything like face value.
- Treat everything else the trace surfaced as a candidate, not a result - a queue of things to test, ordered by whether they are plausible and worth money, not by how striking they looked.
Halo: why the winners look better in every dimension
The fourth mechanism is not statistical at all, and it is the one that contaminates the data before you touch it.
Phil Rosenzweig's account of the halo effect (California Management Review, 2007 - a critical essay re-analysing others' studies, not original empirical research) starts from a much older finding: "First identified by American psychologist Edward Thorndike in 1920, the halo effect describes the basic human tendency to make specific inferences on the basis of a general impression." Applied to company performance, he describes the consequence directly: "When a company is doing well, with rising sales, high profits, and a surging stock price, observers naturally infer that it has a smart strategy, a visionary leader, motivated employees, excellent customer orientation, a vibrant culture, and so on." The inference runs backwards from the outcome: "Rather, company performance creates an overall impression that shapes how we perceive its strategy, leaders, employees, culture, and other elements."
In a closed-won analysis this is not an abstraction, because your CRM is full of human judgement recorded after the fact or alongside it. A rep writes richer notes on a deal that is going well. The qualification reason on a deal that closed reads more favourably than the same circumstances would have read on a deal that stalled. Win/loss debriefs are conducted knowing the outcome. Every one of those fields was recorded by someone who knew, or could sense, which way the deal was going - and each is a field your trace will happily treat as an independent characteristic of the account.
The practical consequence is that a "difference" between wins and losses can be a difference in how they were written up. The fields most exposed are the subjective ones: qualification notes, engagement scoring, sentiment, "champion identified," relationship strength. The fields least exposed are the ones a system stamped without a human in the loop, at a time when the outcome was still unknown - the timestamp of a page view, the campaign a click carried, the firmographics on the account when it was created.
Where a comparison has to lean on a subjective field, the defensible moves are to prefer measurements that predate the outcome, and to have someone code the field without seeing which deals closed. Recommendation Neither is free. Both are cheaper than presenting a halo as a finding.
Why more deals will not settle it
"Come back when we have a year of data" is the standard response to all of this, and it is worth spending a section on why it does not work - because it sounds rigorous, and it postpones the actual fix by a year.
Rosenzweig draws the distinction that answers it, in a passage rebutting a defence of one of the studies he re-analyses:
"Yet such an argument overlooks the critical distinction between noise and bias. If errors are randomly distributed, we call that noise. If we gather sufficient quantities of data, we may be able to detect a signal through the noise. However, if errors are not distributed randomly, but are systematically in one direction or another, then the problem is one of bias - in which case gathering lots of data doesn't help."
That is the whole argument in three sentences. Volume narrows uncertainty. It does not remove bias. Selection, survivorship and halo can each create structured bias whose direction and magnitude depend on the data-generating and recording process - not a fixed push applied to every record, but a distortion built into how the records came to exist. Adding observations drawn from that same process may tighten the interval around a biased estimate without shifting it.
Multiplicity behaves differently again, and the honest version is conditional. More data improves power for a fixed, pre-specified family of comparisons. It creates more opportunities for false discovery only when the analyst also expands what is searched - more dimensions, more cross-tabs, more exclusions tried - which is the normal reflex when a larger dataset arrives.
There is one thing a larger sample genuinely buys, and it is worth naming so this does not read as an argument against ever collecting more: it buys the ability to distinguish effects you have designed a comparison for. That is a real and necessary gain. It is simply a different problem from the four above, and it arrives only after the comparison is built.
The first diagnostic gate - and the confirmation it still needs
Start with what significance does and does not settle. A p-value answers a narrow question under a stated statistical model. If the comparison was selected after searching the data - which is exactly how a backward trace produces candidates - its ordinary interpretation does not account for that selection, or for the wider family of tests the search implicitly ran. And it says nothing at all about selection bias, survivorship or halo, which are properties of how the records came to exist rather than of the test applied to them.
That is a recognised problem with a name. Yoav Benjamini puts it directly: "Selective inference is focusing statistical inference on some findings that turned out to be of interest only after viewing the data. Without taking into consideration how selection affects the inference, the usual statistical guarantees offered by all statistical methods deteriorate. Since selection can take place only when facing many opportunities, the problem is sometimes called the multiplicity problem" (Harvard Data Science Review, 2020). Taylor and Tibshirani frame the same issue in the language of model selection: in their account, most statistical analyses involve some kind of selection - searching through the data for the strongest associations - which is why measuring the strength of the resulting associations is difficult, since the effects of that selection have to be accounted for (PNAS, 2015). That summary is paraphrased rather than quoted here, because the paper sits behind an access barrier this round of checking could not get past, and a paraphrase should not be dressed up in quotation marks it cannot back.
The gate: does the pattern distinguish wins from comparable losses?
A comparison group is the first credibility gate. It tells you whether the pattern seen among winners also distinguishes them from comparable losses - which is worth knowing, and which most backward analyses never establish. Recommendation Five steps, in this order:
- Fix the outcome and the eligible cohort first. Cases are the closed-won opportunities in a mature cohort; the primary controls are comparable closed-lost opportunities from that same cohort. Records that never reached the stage in question belong to a different analysis. The pillar sets out why, and why still-open deals are censored rather than losses.
- Record how the candidate was found - every dimension examined, every cross-tab, every exclusion tried, every lookback rule tested, including the ones that showed nothing. This is the search history, and it is what makes the next step honest.
- Match or stratify on what governed eligibility - cohort period, segment, market, product line, stage. Measuring content-market fit sets out the matching fields this site uses and the reason matched comparison is not optional; that page owns the list and this one does not restate it.
- Measure the same fields on both sides to the same standard, and where you cannot - because losses are recorded more thinly, which is the normal case - say so and account for it rather than quietly comparing a rich record against a sparse one.
- Report the denominator of comparisons from step 2 alongside the result. It is the cheapest honesty measure on this list, and in our experience the one most often skipped.
Why passing the gate is not confirmation
The candidate was selected by inspecting the winners. Testing it against losses drawn from that same cohort uses the data that produced the hypothesis to evaluate it, and writing the claim down at that point does not make it pre-specified - the selection event has already happened. A same-cohort comparison that survives is an exploratory association with better support than it had an hour ago. It is not a confirmed finding, and labelling it one is the single most consequential error available at this stage.
What converts it is data that played no part in the discovery: Recommendation
- Freeze the claim - the dimension, the definitions, the direction, and the minimum difference that would count as meaningful - before looking at anything new.
- Test it on a reserved holdout or a later mature cohort that was not used to find it. Where deal volume is too low to reserve part of the current cohort - which is the constraint most B2B companies are actually under - a later period is the only holdout available, and waiting for it to mature is a real cost worth stating out loud rather than skipping past.
- Where no independent cohort exists, report the pattern descriptively - as an exploratory association, with its uncertainty and the full denominator of comparisons attached - rather than presenting it as confirmation. That is a legitimate output. Presenting the same numbers as a validated finding is not.
A pattern that clears the gate and then repeats out of sample is still an association rather than a demonstrated cause. What it has become is a hypothesis worth spending money to test prospectively, which is a genuinely different thing from a hypothesis worth spending money to act on.
When the honest answer is "we cannot tell"
Sometimes the comparison cannot be made, and saying so is a result rather than a failure.
Small deal counts are the usual reason, and it is worth being precise about why rather than reaching for a number.
NIST's engineering statistics handbook works an example that shows how quickly sample requirements grow even for a question simpler than the one a trace asks: testing a single observed proportion against a fixed 10% benchmark and detecting a rise to 20% requires roughly 102 observations, or 112 with its continuity correction, under a one-sided test at a 5% significance level with 90% power (NIST/SEMATECH e-Handbook of Statistical Methods, section 7.2.4.2; a living document with no per-section revision date, reviewed here on September 4, 2026, which attributes the derivation to Fleiss, Levin and Paik).
That number cannot be carried across to a two-channel comparison, and the mismatch is worth naming because it is an easy one to make. NIST is sizing a one-sample test against a fixed benchmark. Comparing conversion between two channels is a two-proportion problem, and its denominators are not your win and loss counts - they are the opportunity and win counts within each channel. Eleven wins and fourteen losses are outcome groups, not channel denominators. To know what your own data can separate, calculate power from the opportunity and win counts in each channel, and state every assumption: allocation between groups, significance level, power, test direction, and the minimum difference you would act on.
What follows from a small cohort is low power and wide uncertainty, not the impossibility of testing - an exact test can be run on a small table. The precise statement is that the design lacks the resolution to separate the alternatives, not that the pattern is absent, and the difference between those two sentences is the difference between an honest report and a misleading one. The arithmetic for your own numbers, and how it should set where you start, is on the pillar's starting-point ladder and in the sub-pillar that owns that decision.
Two other honest stopping points. Where the losses were recorded so thinly that the comparison would be between a full record and an empty one, the finding is about your data collection rather than your funnel. And where a pattern rests entirely on subjective fields written after the outcome was known, it should be reported as what it is - a description of how won deals get written up.
None of these is a wasted analysis. Each one tells you something specific about what to fix before the next attempt, which is more than a confidently presented artefact would have done.
What a surviving pattern has earned
A pattern that came through a defined cohort, a comparison group, matched strata, symmetric measurement and a stated denominator of comparisons has earned exactly one thing: a place at the front of the queue. If it then repeats on a cohort that played no part in finding it, it has earned a stronger claim than almost anything else in a marketing report. It is still not evidence that the thing you noticed caused the wins, and the four mechanisms on this page are the reasons why not - each of them can survive a well-built comparison in weakened form.
What converts an association into something you can act on is a prediction, written down before the period it applies to, with a comparison built in. That is the subject of the next page in this cluster, and it is where the campaign you were about to launch belongs: not as a rollout, but as the instrument of the test.
Sitting on a finding you are not sure about? The quickest check is the denominator: write down every dimension the trace examined, not just the one that stood out. If that list is long and the comparison group is missing, you have a candidate rather than a result - and that is worth knowing before the deck, not after.
Sources
Sources. Selection and survivorship: Miguel A. Hernán, Sonia Hernández-Díaz and James M. Robins, "A Structural Approach to Selection Bias," Epidemiology 15(5), 2004, pp. 615-625 - a theoretical paper using causal diagrams; the authors note their structural definitions "might not coincide perfectly with the traditional, often discipline-specific, terminologies." Jerker Denrell, "Vicarious Learning, Undersampling of Failure, and the Myths of Management," Organization Science 14(3), 2003, pp. 227-243 - an analytical paper about organizational survival, not an empirical study of firms; its application to sales cohorts here is an inference from a shared mechanism. Jerker Denrell and Balázs Kovács, "Selective Sampling of Empirical Settings in Organizational Studies," Administrative Science Quarterly 53(1), 2008, pp. 109-144 - a simulation study, and the source for the point that the bias is correctable. Stephen J. Brown, William Goetzmann, Roger G. Ibbotson and Stephen A. Ross, "Survivorship Bias in Performance Studies," Review of Financial Studies 5(4), 1992, pp. 553-580 - growth equity mutual funds, a domain with far more complete data than a CRM; its finding is "the appearance of predictability," which is not the same claim as spuriousness or false persistence and is not blended with either here. Abraham Wald, A Reprint of "A Method of Estimating Plane Vulnerability Based on Damage of Survivors," Center for Naval Analyses report CRC 432, July 1980, also available at DTIC accession AD-A091 073, reprinting memoranda written for the Statistical Research Group at Columbia University during the war and printed there as Parts I-VII (secondary accounts describe eight memoranda; the reprint's own contents list seven, so no count is asserted here) - cited for the formal impossibility result only; the mathematical notation renders inconsistently across copies and is described rather than quoted, and the familiar armour-placement story is a later retelling rather than Wald's wording, as set out in Bill Casselman's "The Legend of Abraham Wald" for the American Mathematical Society Feature Column, accessed September 4, 2026. Multiplicity: Joseph P. Simmons, Leif D. Nelson and Uri Simonsohn, "False-Positive Psychology," Psychological Science 22(11), 2011, pp. 1359-1366 - the rates quoted are from computer simulations of 15,000 samples in a two-condition design with 20 observations per cell, reported at p < .05 alongside p < .1 and p < .01, and are not observed rates in any literature or dataset. John P. A. Ioannidis, "Why Most Published Research Findings Are False," PLoS Medicine 2(8), 2005, e124 - an essay presenting an analytical and simulation framework, not a measurement of any body of published work; a 2022 correction (PLoS Medicine 19(8), e1004085) fixed a missing set of parentheses in one equation in Table 2 and did not alter the argument. Yoav Benjamini and Yosef Hochberg, "Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing," Journal of the Royal Statistical Society: Series B (Methodological) 57(1), 1995, pp. 289-300 - the false-discovery-rate definition is described rather than quoted, because the equation did not reproduce identically across two clean reads; the theorem quoted is proved for independent test statistics, a condition stated here every time the paper is cited. Selective inference: Yoav Benjamini, "Selective Inference: The Silent Killer of Replicability," Harvard Data Science Review 2(4), 2020 - one of a twelve-article special feature on reproducibility and replicability; accessed September 4, 2026, and note the article carries numbered revisions on its platform. Jonathan Taylor and Robert J. Tibshirani, "Statistical learning and selective inference," Proceedings of the National Academy of Sciences 112(25), 2015, pp. 7629-7634 - paraphrased rather than quoted, because the full text was not reachable from our tooling and a paraphrase should not be dressed up as a direct quote. Yoav Benjamini and Daniel Yekutieli, "The control of the false discovery rate in multiple testing under dependency," The Annals of Statistics 29(4), 2001, pp. 1165-1188 - quoted from the abstract; the theorems and the PRDS definition are described rather than quoted because the mathematical notation did not reproduce identically across reads, and both results establish control at a level bounded by the proportion of true null hypotheses rather than exactly at q. Halo: Phil Rosenzweig, "Misunderstanding the Nature of Company Performance: The Halo Effect and other Business Delusions," California Management Review 49(4), Summer 2007, pp. 6-20 - a critical essay re-analysing other researchers' studies; it presents no original data, and the noise-versus-bias passage is quoted from the start of its argument rather than from the closing clause, which cannot stand alone. Sample size: NIST/SEMATECH e-Handbook of Statistical Methods, section 7.2.4.2, "Sample sizes required" - sample size for a hypothesis test about a proportion, not for margin-of-error estimation; the worked example is one-sided at a 5% significance level with 90% power, the handbook attributes the derivation to Fleiss, Levin and Paik, and it is a living document with no per-section revision date, reviewed on September 4, 2026.
About the author
Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.
This page is part of the reverse funnel marketing cluster. The rest of the sub-pillars are in production and will be linked here as they publish.