Two labels mark evidence boundaries on this page. Documented marks a platform definition or behaviour directly supported by the linked provider documentation. Recommendation means a workflow, threshold or decision rule that CoreAEX prescribes and no source specifies. Published research is attributed in the prose with its venue, design and scope, and is not tagged - it is evidence about a method, not documentation of how a product behaves. This page cites no platform documentation, so no Documented tag appears on it; that is expected rather than an omission. Untagged text is ordinary explanation or a conclusion following from something already labelled.
You want to start at closed-won deals, because that is the outcome you care about. Then you count: fourteen wins in the period, four of which closed last month, and an unknown number of deals still open that might yet go either way.
Whether that is enough is not a matter of preference or nerve. It is a property of your data, and two constraints decide it - cohort maturity and sufficiency. Both require rules you set before you look at the results, and neither of them yields a threshold you can memorise. What they yield is an honest answer to a narrower question: given what you have, which starting point can carry a conclusion, and how strong is the conclusion allowed to be?
Two things this page will not do, because both are common shortcuts and both are wrong. It will not treat a lower rung as a way of escaping the problem of unfinished deals - a lower rung changes what you are measuring. And it will not tell you that a given number of wins is or is not enough, because a win count on its own cannot answer that.
The operating rule from the pillar stands: never begin farther from the outcome than your data forces you to. This page is how you work out where that is.
The four rungs, and what each one costs
The ladder runs from the business outcome itself to progressively more distant proxies for it. Every step upward buys you volume and timeliness, and replaces the outcome you care about with something that stands in for it.
| Rung | Use it when | What it costs |
|---|---|---|
| Closed-won opportunities, with a named value measure | You have enough mature, reliable outcome data | Least proxy loss: the recorded outcome is the business outcome. You still need a comparison group, uncertainty intervals and controls for selection and confounding before you generalise. The cost is volume and waiting |
| SQLs or qualified opportunities | Wins are too few, or too many deals are still open | You avoid waiting for final revenue by replacing it with a proxy outcome - qualification, or opportunity creation. Later win or loss stays unknown, and must not be inferred from stage attainment |
| Genuinely defined MQLs | Reliable opportunity data does not exist | The measurable outcome is a marketing-defined qualification event. Useful for diagnosing acquisition patterns, but its relationship to revenue has to be estimated separately, and it can differ by segment, period and scoring rules |
| Leads | Diagnostic use only | It never observes an outcome. It can describe where volume comes from; it cannot distinguish the volume that becomes revenue |
One correction before the ladder is used, because it changes what the upper rungs are for. Moving up to a stage outcome does not relocate the problem of unfinished deals. If the outcome you are measuring is "reached SQL," then an opportunity that reached SQL has had its event observed; nothing about it is unresolved for that question. What remains unknown is a different outcome - whether it will eventually be won - and that is an open business question, not a statistical artefact of the analysis you are running. The trade is timeliness for a proxy, and it is a real trade with a real cost. It is just not the cost of censoring.
The definitions of MQL and SQL, and the reasons an MQL count is not comparable between companies, belong to other pages in this cluster - measuring content-market fit owns the definitions, and what your lifecycle and attribution fields record owns why the label is configured rather than measured. What this page owns is the choice between them.
Constraint one: is the cohort mature?
Deals still open have not had an outcome. They are not wins, they are not losses, and putting them in either group misrepresents them. The statistical name for their situation is right censoring, and it has a condition attached that matters more here than the name does.
Clark, Bradburn, Love and Altman set out the concept in a peer-reviewed tutorial on survival analysis (British Journal of Cancer, 2003). Censoring arises, among other ways, when "a patient has not (yet) experienced the relevant outcome, such as relapse or death, by the time of the close of the study" - and "this situation is often called right censoring." The clinical vocabulary transfers: an opportunity still open at the end of your observation window occupies the same position in the analysis - a record whose outcome exists but has not been observed yet.
The condition is the part to hold on to:
"Standard methods used to analyse survival data with censored observations are valid only if the censoring is 'noninformative'. In practical terms, this means that censoring carries no prognostic information about subsequent survival experience; in other words, those who are censored because of loss to follow-up at a given point in time should be as likely to have a subsequent event as those individuals who remain in the study."
Read that condition carefully, because it is about why observation ended - not about why the deal has not closed. Notice that the passage is phrased around loss to follow-up: the question it asks is whether the records that stopped being observed differ, in their prospects, from the ones still being watched.
The distinction matters, and it is the one most easily lost. At a fixed administrative cutoff, open deals are censored because their event times extend beyond the observation window. Why they remain open may well predict their eventual outcome or their time to close - but that alone does not make the censoring mechanism informative. Informative censoring becomes a concern when observation ends, records disappear, or follow-up differs for reasons associated with the future outcome: when likely losses stop being updated and quietly fall out of the working set, when historical records are missing selectively, when a segment or a legacy system is left out of the export, or when local habit marks some stalled deals closed-lost and leaves others open indefinitely. Which of those describes your data is a judgement about your own pipeline rather than a finding in the paper, and it is the reason cohort maturity is a design decision rather than a formality.
The mistake this leads to is filtering rather than reporting. The survival methods that handle censoring keep the censored records and model them. Deleting every open opportunity is not a way of solving censoring; it is a sampling rule, and if the open records differ systematically from the closed ones it introduces the selection problem it appears to avoid. Waiting longer genuinely reduces the unresolved share, but a mature closed-only cohort is not automatically an unbiased one.
So the rule has two halves: fix the cohort before you look, and then report three outcome states rather than two. Recommendation Set the cohort-entry window and the follow-up horizon in advance and in writing, anchored to your own median and long-tail sales-cycle length rather than to a quarter boundary. At the cutoff, report won, lost and still-open records separately. Never recode an open opportunity as a loss and never remove one silently. If the analysis is descriptive, disclose the unresolved share and check whether the conclusion survives the two extreme assumptions - every open deal a win, then every open deal a loss; if the answer flips, you do not yet have a finding. If it is genuinely inferential about time to outcome, use survival or competing-risk methods and state their assumptions. And a cohort chosen after seeing which window produces the tidier result is a selected cohort, whatever the window length; the page on telling a pattern from an artefact explains what that does to everything downstream.
Constraint two: what can your numbers separate?
The second constraint is the one where confident-sounding numbers get invented, so this section is deliberately careful about what each figure applies to.
Start with the quantity that makes the question answerable. Howard Bloom defines it this way (MDRC working paper, 2006 - a methods working paper, not a peer-reviewed article):
"Intuitively, a minimum detectable effect is the smallest true treatment effect that a research design can detect with confidence. Formally, it is the smallest true treatment effect that has a specified level of statistical power for a particular level of statistical significance, given a specific statistical test."
That definition contains the whole discipline of this section. A minimum detectable effect is not a property of your data alone; it is a property of a design - a stated test, a stated significance level, a stated power. Change any of those and the number changes. Which is why "how many deals do I need?" has no answer until you have said what you want to detect and how confident you intend to be.
The one-sample case: your rate against a fixed benchmark
The simplest version of the question is whether one observed proportion differs from a fixed reference value. NIST's engineering statistics handbook works an example: a manager needs to detect any change above 0.10 in a proportion currently running at approximately 10%, wants a one-sided test, accepts a 5% significance level, and is willing to take a 10% risk of failing to detect a change of that size. The required sample is roughly 102, or 112 with the continuity correction (NIST/SEMATECH e-Handbook of Statistical Methods, section 7.2.4.2, "Sample sizes required"; a living document with no per-section revision date, reviewed here on September 4, 2026, which attributes the derivation to Fleiss, Levin and Paik).
Note every one of those conditions. One proportion. A fixed benchmark. One-sided. Alpha 5%. Power 90%. A ten-point difference to detect. Drop any of them and the number is no longer 102.
The two-channel case is a different calculation
The question people actually want to ask is not the one above. It is: does channel A convert better than channel B? That compares two observed proportions against each other, which is a different design with its own sample-size arithmetic. Recommendation Do not reuse a one-sample figure for it - the arithmetic does not transfer, and the substitution is easy to make because both problems look like "how many do I need."
Two things go wrong here, and the second is more damaging than the first.
The denominators are not your win and loss counts. Fourteen wins and eleven losses are outcome groups. A channel comparison needs, within each channel, the number of opportunities that entered and the number that were won. A channel that produced three opportunities and two wins has an observed opportunity-to-win rate of 66.7% - that is simply the arithmetic, and there is nothing wrong with saying so. What is wrong is carrying it forward as though it were stable. Report the numerator, the denominator and an interval, and let the reader see that the estimate rests on three observations.
And the parameters have to be stated, not assumed. A two-proportion calculation requires the two expected proportions or the minimum difference worth detecting, the allocation between the groups, the significance level, the power, and whether the test is one- or two-sided. A sample size quoted without those is not a result.
For a sense of scale: a peer-reviewed methods paper documenting the standard normal-approximation formula works an example comparing two proportions of 90% and 80%, at a 5% significance level with 80% power and equal group sizes, and reports that 219 observations are needed in each group, 438 in total, with a continuity correction applied (Geoffrey T. Fosgate, Journal of Veterinary Diagnostic Investigation, 2009). The worked case is a diagnostic-testing question in veterinary medicine; what transfers is the formula and the parameter list, not the subject matter.
Now hold that against a realistic B2B cohort - carefully, because the obvious conclusion is also wrong. The obvious conclusion is that fourteen wins across all channels settles it: nowhere near 219 a group, so nothing can be separated. But fourteen wins is not a sample size. It is a numerator, and the same numerator can carry very different evidence depending on the denominators underneath it.
- Channel A wins 10 of its 10 opportunities; channel B wins 4 of its 18. Fourteen wins in total, and Fisher's exact test on that table returns a two-sided p of about 0.0002.
- Channel A wins 7 of its 10; channel B wins 7 of its 18. The same fourteen wins, and the same test returns about 0.24.
Fisher's exact test is one of the two methods NIST's handbook documents for comparing two proportions (section 7.3.3, alongside the large-sample z-statistic) - and that section documents the test only. It carries no sample-size formula, which is part of why the one-sample figure gets borrowed for a job it cannot do.
So the defensible statement is not that fourteen wins cannot separate two channels. It is that fourteen wins on their own cannot tell you either way. Recommendation Before making any claim about whether your counts can separate two channels, write down five things: the opportunities that entered each channel, the allocation between them, the proportions you expect or observe, the difference that would change a decision, and the test you intend to run. The 219-per-group figure is what one specific design costs - 90% against 80%, 5% alpha, 80% power, equal groups - and not a general B2B minimum. Separating two groups that both convert well, by ten points, is an expensive thing to ask for. Separating 100% from 22% is not.
Tests and intervals are also not the same instrument, and small samples do not disable either one automatically. An exact test can control the type-I error rate in a small table under its assumptions, and an extreme table can still be distinguishable, as the first example above shows. What small samples cost you is power against modest differences. So pair the test with an appropriate interval, and do not read either significance or inconclusiveness off the sample size alone.
What a small cohort can still support
Two pieces of statistical practice are worth knowing here, because they define the edges of what small numbers permit.
When nothing happened at all, there is a usable bound. Hanley and Lippman-Hand's rule of three (JAMA, 1983) states that "if none of n patients shows the event about which we are concerned, we can be 95% confident that the chance of this event is at most three in n (ie, 3/n)." So if none of your twenty enterprise opportunities came through a particular channel, you can say with 95% confidence that the underlying rate is at most about 15% - which is a real, defensible statement, and a good deal more useful than "we saw none."
Two scope conditions travel with it, and both matter. It is a 95% one-sided upper bound given a zero numerator - not a general small-sample rule and not applicable when you observed one or more events. And the authors note their own approximation's limits: "when n is larger than 30, the rule of three agrees with the exact calculation to the nearest percentage point; below 30, it slightly overestimates the risk, but then the maximum risk [ie, 26% if n=10] is becoming so high that the difference between it and that suggested by the rule [3/10=30%] is not usually worth quibbling about."
A third condition is about scope rather than arithmetic, and it is the one that gets dropped first. The bound describes the process and population those twenty observations came from, treating them as independent observations of a stable underlying success rate. It does not travel to a different segment, a different period, or a campaign whose selection and conversion behaviour differ - "at most 15%" is a statement about that channel in that window, not a standing property of the channel.
When something did happen, the interval you put around it matters more than the point estimate. Brown, Cai and DasGupta examined how confidence intervals for a proportion behave (Statistical Science, 2001, published with invited discussion at pp. 101-133) and are blunt about the textbook method: "it is widely recognized that the actual coverage probability of the standard interval is poor for p near 0 or 1," and "the popular prescriptions the standard interval comes with are defective in several respects and are not to be trusted."
Their recommendation for the sample sizes a B2B cohort actually has is explicit: "for small n (40 or less), we recommend that either the Wilson or the Jeffreys prior interval should be used." For larger samples they find the Wilson, Jeffreys and Agresti-Coull intervals broadly similar and favour Agresti-Coull on simplicity. Recommendation The practical consequence for a fourteen-deal cohort is worth stating precisely, because it is easy to get backwards. Check what your spreadsheet is actually computing - the default is tool-dependent - and if it is the standard Wald interval, replace it with Wilson or Jeffreys. Then expect the bounds to move rather than simply to widen. Around the middle they can tighten: seven wins from fourteen opportunities gives a Wald interval about 0.52 wide and a Wilson interval about 0.46. Near the ends they widen: one win from fourteen gives about 0.27 against about 0.30. And the Wilson interval stays inside 0 and 1, where the Wald interval on two wins from three opportunities runs past 1 altogether. The reason to switch is coverage behaviour, not a guarantee of conservatism.
Why waiting for more deals is not automatically the answer
"We will run this again next year with a full cohort" is the natural response to everything above, and it is right about one thing and wrong about another.
It is right that volume buys resolution. Every constraint in the two sections above eases with more observations: intervals generally narrow, minimum detectable effects shrink under otherwise unchanged design assumptions, and the analysis gains power to detect a difference if one exists. If the problem is genuinely that your design lacks power, more mature deals is the fix.
It is wrong if the problem is bias rather than noise, and the two are easy to confuse from the inside. Selection, survivorship and halo effects are structural: they arise from how the records came to exist, so drawing more observations from the same process tightens the interval around an estimate without shifting it toward the truth. The page on patterns and artefacts works through the four mechanisms and the reason more data addresses none of them directly.
The obvious shortcut for telling them apart does not work, and it is worth saying why. The shortcut is to ask, in the meeting, whether more deals would change the answer or merely show the same pattern more clearly - and to treat the second reply as a symptom of bias. It is not. Showing the same pattern more clearly is also what more observations do when the pattern is real: reducing random error is the job additional observations perform. The question cannot separate a genuine signal from a structural one, so it should not be relied on as though it could.
Diagnose the two separately instead. Recommendation For random error, use the instruments above: denominators, intervals, and the difference the design could have detected. For bias there is no equivalent single number, so work through the structure of the data. Write down the rule that decided which records entered the analysis. Compare the included records with the excluded ones on whatever you can still observe about both. Look at which fields are missing, and for which deals, rather than only at the rows that are complete. Check where each field's value came from and who or what set it. Then, where the data allows it, run a sensitivity check across the assumptions you had to make, a negative control - something the mechanism you are proposing should not affect - or a holdout period, and see whether the pattern reproduces. If the collection looks sound and the uncertainty is simply wide, waiting for more mature records is the fix. If the collection itself is doing the selecting, more of it will not help.
The decision, in order
Six steps, and the order matters because each one can end the process. Recommendation
- Write down the observation window that makes a cohort mature, anchored to your own sales-cycle distribution, before looking at any results. Then identify which cohorts clear it.
- Count won, lost and still-open records separately for the cohorts that clear the window, and record the unresolved share. A wins-only count answers a different and lesser question; an open record folded into either group answers a false one.
- State the comparison you want to make and the difference that would change a decision. Not "which channel is best" but "is channel A's opportunity-to-win rate at least X points higher than channel B's, and would that change where we spend?" Without the second half, no sample size can be calculated.
- Work out whether your denominators can detect that difference - opportunities entered within each group, not wins against losses - using the arithmetic for the design you are actually running, one-sample against a benchmark or two-sample between groups, with allocation, significance level, power and test direction all stated.
- If they can, start at closed-won. If they cannot, decide deliberately between three options rather than defaulting: move up a rung and accept a proxy outcome, knowing it answers a different question; widen the window and wait; or keep the closed-won cohort and report it descriptively, with denominators, intervals and the unresolved share all shown.
- Write down which rung you chose and why, in the same document as the analysis. The reason is not bureaucracy: six months later, the choice will look like a finding unless the constraint that produced it is recorded next to it.
Moving up a rung is a legitimate response and not a defeat, but it is not free, and the page that owns what each substitution costs is the pillar's ladder in summary and the sub-pillar on proxy outcomes in full.
What each rung lets you say
The point of the whole exercise is the sentence you get to write at the end, so it is worth setting out what each starting point actually licenses.
| Starting point | What you can say | What you cannot |
|---|---|---|
| Won and lost outcomes with usable denominators, unresolved share small and disclosed | These characteristics were associated with winning rather than losing in this cohort, with this interval and these stated design assumptions | That they caused the wins, that the pattern will repeat, or that a closed-only cohort represents the population it was drawn from |
| Won and lost outcomes, small counts | Here is what we observed, here is the interval around it, and here is the difference the design could have detected | That one channel outperformed another when the interval covers both possibilities - though an extreme table may still separate, so state the difference you could have detected rather than the count |
| Won and lost outcomes, but a material share still open | The three outcome states, reported separately, with a fixed horizon and a sensitivity check across the unresolved records - or time-to-event methods if the question is about timing | Drop the open deals silently, recode them as losses, or describe what remains as unbiased because it is "mature" |
| SQL or qualified-opportunity cohort | These characteristics were associated with reaching that stage | Anything about revenue - the outcome measured is a proxy, and later win or loss cannot be inferred from stage attainment |
| MQL cohort | These characteristics were associated with being labelled qualified by our own team, under the scoring rules in force at the time | That the finding is about buyers rather than partly about your qualification rules, or that the label means the same thing across segments and periods |
| Lead-level only | Volume, mix, source coverage and data quality: volume came disproportionately from here | That lead patterns are pipeline or revenue drivers, or anything distinguishing volume that converts from volume that does not |
Every row in that table is a real conclusion that some analysis is entitled to reach. The failure this page exists to prevent is not starting in the wrong place - it is starting in the right place and then writing the sentence from a row above it.
Not sure which rung your data supports? Start at step 3: write down the specific difference that would change a decision, in percentage points. If that number does not yet exist, no sample-size question has an answer - and agreeing what result would actually move the budget becomes the first task rather than the last.
Sources
Sources. Censoring and cohort maturity: T. G. Clark, M. J. Bradburn, S. B. Love and D. G. Altman, "Survival Analysis Part I: Basic Concepts and First Analyses," British Journal of Cancer 89(2), 2003, pp. 232-238 - a peer-reviewed tutorial in a clinical context; the application to open sales opportunities, and the account of which pipeline conditions would breach the noninformative-censoring condition, are inferences about the setting rather than findings in the paper. The page follows the paper in locating that condition at loss to follow-up - why observation ended - rather than at the length of an unresolved event time; a fixed administrative cutoff is not treated here as informative censoring on its own, however long the open deals have been open. The paper supports the definition of censored observations and the methods that retain them; it is deliberately not used here to justify dropping open opportunities from a cohort. The underlying estimator originates with E. L. Kaplan and Paul Meier, "Nonparametric Estimation from Incomplete Observations," Journal of the American Statistical Association 53(282), 1958, pp. 457-481, which uses the term "loss" rather than "right censoring" - the modern term postdates it and is not attributed to them here. Minimum detectable effect: Howard S. Bloom, "The Core Analytics of Randomized Experiments for Social Research," MDRC Working Papers on Research Methodology, 2006 - an MDRC working paper, not a peer-reviewed journal article; used for the design-specific definition of a minimum detectable effect, which is a property of a stated design rather than an intrinsic property of a dataset. Sample size, one-sample: NIST/SEMATECH e-Handbook of Statistical Methods, section 7.2.4.2, "Sample sizes required" - a test of one observed proportion against a fixed benchmark, one-sided in the worked example, reviewed September 4, 2026; the handbook is a living document with no per-section revision date and attributes the derivation to Fleiss, Levin and Paik. The NIST handbook documents no sample-size section for comparing two proportions - its chapter 7 lists "Sample sizes required" only under one-process sections, and its two-process material on proportions (section 7.3.3, which documents the large-sample z-statistic and Fisher's exact test) covers the test rather than the sample size - so it is deliberately not cited for the two-channel case. The two illustrative tables used on this page, 10 of 10 against 4 of 18 and 7 of 10 against 7 of 18, are arithmetic rather than sourced findings: the two-sided p-values quoted, approximately 0.0002 and 0.24, are Fisher's exact test computed on those tables and are reproducible by any implementation of it. Sample size, two-sample: Geoffrey T. Fosgate, "Practical Sample Size Calculations for Surveillance and Diagnostic Investigations," Journal of Veterinary Diagnostic Investigation 21(1), 2009, pp. 3-14 - a peer-reviewed methods paper documenting the standard normal-approximation formula for comparing two proportions, assuming equal group sizes and a two-sided alpha; the worked example concerns diagnostic-test sensitivity in veterinary medicine, and only the formula and its parameter list are carried across. The 219-per-group figure supports that specific 90%-against-80% design and is not presented here as a general B2B minimum. Small samples: James A. Hanley and Abby Lippman-Hand, "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators," JAMA 1983;249(13):1743-1745 - the rule of three is a 95% one-sided upper bound applicable to a zero numerator, not a general small-sample rule, and the authors' own note on its behaviour below n=30 is quoted alongside it. The bound assumes independent observations of a stable underlying rate and describes only the process and population those observations came from; the scope caveat on the page states that and is not drawn from the paper. Lawrence D. Brown, T. Tony Cai and Anirban DasGupta, "Interval Estimation for a Binomial Proportion," Statistical Science 16(2), 2001 - published with invited discussion and rejoinder across pp. 101-133, with the authors' own article ending earlier in that range; their small-sample recommendation is quoted directly and their guidance for larger samples is paraphrased, because its exact bracketing did not reproduce identically across reads. Their case for Wilson and Jeffreys rests on coverage behaviour, not on interval width; the interval-width comparisons on this page are arithmetic worked here, not findings taken from the paper, and the Wald and Wilson intervals quoted are the standard 95% forms.
About the author
Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.
This page is part of the reverse funnel marketing cluster. The rest of the sub-pillars are in production and will be linked here as they publish.