Two labels mark evidence boundaries on this page. Documented marks a platform definition or behaviour directly supported by the linked provider documentation. Recommendation means a workflow, threshold or decision rule that CoreAEX prescribes and no source specifies. Published research is attributed in the prose with its venue, design and scope, and is not tagged - it is evidence about a method, not documentation of how a product behaves. Untagged text is ordinary explanation or a conclusion following from something already labelled.

Eleven wins in eighteen months. A board meeting in three weeks. Someone sensible has already said the sensible thing - look at MQLs instead, you have hundreds of those - and they are not wrong to suggest it.

What the suggestion needs alongside it is a statement of what the substitution costs, because the cost is not a vaguer version of the same answer. It is a different answer to a different question, and the difference has to be carried through into the sentence you eventually write. Where to start a backward analysis owns the decision about which rung your data supports. This page is about what happens after you have made it: what you are now measuring, what your comparison group has to become, and which claims quietly stop being available.

What actually changes, and two things that do not

Moving up a rung replaces the outcome with a proxy for it. That is the whole of the change, and stating it plainly is more useful than it sounds, because two other explanations are available and both are wrong.

It is not that you have escaped incomplete observation

The first wrong explanation is that upstream analysis gets you out of the problem the unfinished deals were causing - that they were the obstacle, and moving to an earlier stage steps around them.

What the earlier endpoint changes is which records have had their event observed. If your outcome is "reached SQL," then a record that reached SQL has had its event observed, and nothing about it is unresolved for that question. What remains unknown - whether the deal will eventually be won - is a different outcome, and an open business question rather than a property of the analysis you are running. Moving upstream often shortens the follow-up you need. It does not remove incomplete observation, and whether you end up with fewer unresolved records depends on the design you choose rather than on the endpoint alone.

So say which design you are running before you classify a single record, because the two treat an incomplete record differently and mixing them is how a record ends up in the wrong group.

  • A fixed-horizon binary analysis asks whether each record reached the endpoint within a prespecified observation horizon. Include only records that have completed that horizon. A record still inside it is unresolved - it is not a "did not reach SQL" control, and putting it there manufactures negatives out of records that have not finished.
  • A time-to-event analysis asks how long records take to reach the endpoint. Here a record whose event has not been observed when follow-up ends is right-censored, in the sense Clark, Bradburn, Love and Altman set out in their peer-reviewed survival-analysis tutorial (British Journal of Cancer, 2003) - and censored records are retained and modelled rather than dropped or recoded.

Recommendation Write the design choice at the top of the analysis document, in one sentence, before the cohort query. Censoring is a time-to-event concept; if you are running a fixed-horizon comparison, the word does not apply to your incomplete records and reaching for it will lead you to treat them as though a survival method were handling them. The cohort-maturity rules apply to whatever endpoint you have chosen - fix the window in advance, report the outcome states separately, and never recode an unresolved record as a negative one.

So upstream analysis does buy you timeliness, genuinely. It buys it by measuring something else.

It is not that you have solved the small-sample problem

The second wrong explanation is arithmetic, and it is the more expensive of the two. Eleven wins felt too few; four hundred MQLs feels like plenty; the sample-size problem therefore feels solved.

It is not solved, because you have changed what is being counted rather than getting more of the thing you were counting. The four hundred are MQLs, and any comparison you run on them needs the denominators for that question: within each group you are comparing, how many records were eligible to reach the MQL stage, and how many did. If your real interest is still revenue, the eleven wins have not become four hundred wins. They are still eleven, and the sample-size arithmetic has not moved.

Moving upstream can genuinely improve precision - for the upstream question. Recommendation Before you move, write down the comparison you intend to run at the new rung, with its denominators, and check that it is a question worth answering rather than the old question wearing a new label. If the answer you need is about revenue and the only defensible analysis is about qualification, that is worth knowing before you build it, not after.

The comparison group moves with the outcome

This is the part that gets missed, and it invalidates more upstream analyses than any other single thing.

The rule from the pillar is one risk set per outcome: fix the outcome first, then the cohort that was eligible to reach it, then the controls. That rule does not stop applying when you move up a rung. It follows you, and it changes your comparison group underneath you if you are not watching.

Where the outcome moves, the eligible cohort and the controls move with it
If your outcome is…The cases areThe controls areWhat is not a control
Closed-wonWon opportunities in a mature cohortComparable closed-lost opportunities from the same eligible cohortDisqualified leads and records that never became opportunities - they were never at risk of this outcome
Reached the defined SQL endpointRecords whose stage history shows they entered the defined SQL stage inside the windowEligible records that completed the prespecified observation horizon without the eventClosed-lost status is neither a case rule nor a control rule. A closed-lost record is a case only when its stage history shows it reached the defined SQL endpoint inside the window; otherwise classify it from the recorded SQL event, not from the later deal result
Reached MQLRecords whose history shows the MQL label was applied inside the windowEligible records that completed the prespecified observation horizon without the eventAnything that entered your database after the window, or through a route that could not produce the label - and any record still inside the horizon, which is unresolved rather than negative

Look at the second row for a moment, because it is counter-intuitive enough to be worth sitting with. When the outcome is "reached SQL," a record whose history confirms SQL entry remains an SQL case even if a later opportunity was lost. It did the thing you are measuring. Sweeping it into the control group because it did not become revenue imports the closed-won question into an analysis that is not about closed-won, and produces a comparison that answers neither.

But notice what that row does not say, because the shortcut is tempting in the other direction too. It does not say that closed-lost opportunities are automatically SQL cases. Lifecycle stage is a field with several entry routes, and HubSpot documents Sales Qualified Lead and Opportunity as separate stages, so a record can carry a deal association without its history showing that it entered SQL under the definition in force for your cohort. Classification comes from the recorded stage event, not from where the record ended up. If your stage history cannot support that, you have found a data problem rather than a cohort - and what your lifecycle stages actually record is where that gets diagnosed.

Recommendation Write the outcome definition at the top of the analysis document, before the query, and check every group assignment against the recorded event rather than against your intuition about which records are the good ones. The intuition is trained on revenue and it will keep pulling the analysis back there.

The rest of the procedure - the joins, the exposure window, the matching discipline, the capture-quality audit - is unchanged, and the backward-trace procedure owns all of it. What changes at each rung is the outcome, the cohort, the controls, and the sentence at the end.

Choose one upstream endpoint: SQL or qualified opportunity

This is the shortest step down and the one with the least loss, which is why it is worth being precise - starting with the fact that the heading contains an "or" rather than a slash.

Choose one event and analyse that one. Recommendation Use SQL entry when the question concerns sales qualification. Use qualified-opportunity or deal-stage entry when the question concerns the creation of qualified pipeline. Combine the two only if a CRM audit shows they identify the same event, with the same timestamp and the same eligibility rule, across the whole cohort window - and if that audit has not been done, they are two endpoints rather than one.

HubSpot's documentation is a good illustration of why the distinction is not pedantry: it defines a Sales Qualified Lead as "a contact or company that your sales team has qualified as a potential customer" and an Opportunity separately as "a contact or company that is associated with a deal." Documented Those are different events with different triggers, and other CRMs and internal processes use still different objects, stages and gates. Merging them changes both the endpoint and the risk set.

Read the SQL definition, too, for what it does not contain: no universal qualification criteria. Your organisation defines and applies the standard, either by hand or through configured automation, and HubSpot records the resulting lifecycle-stage value and its history rather than independently measuring whether qualification occurred. That is a statement about where the judgement lives, not about what the platform is capable of documenting. What your lifecycle stages and attribution fields actually record owns that audit in full.

What a stage-attainment cohort supports, and what it does not

What a cohort at either endpoint supports. A defensible observational finding about what preceded the event: which channels, content and account characteristics were associated with reaching the defined stage rather than stalling before it. The supported form of that sentence is narrower than it feels - "among eligible records in this cohort, channel A was associated with a higher rate of reaching the defined SQL stage than channel B," with the denominator, the effect estimate and the uncertainty reported alongside it. For a marketing team accountable for qualified pipeline, that is not a consolation prize; it is a direct measure of stage attainment under the definition in force.

What it does not support, and this is the boundary that erodes first. Reaching a stage tells you nothing about the outcome beyond it. A channel associated with a higher qualification rate may be associated with a lower win rate, and this analysis cannot see that - it stopped measuring before the point where the difference would appear. Recommendation So never carry a stage-attainment finding into a revenue sentence, and watch the verb as closely as the noun. "Was associated with a higher rate of reaching qualification" is the form the evidence supports. "Produced," "generated" and "drove" all assert that the channel made it happen, which an observational cohort has not shown; and "pipeline" belongs to an explicitly defined opportunity or deal endpoint, not to a contact-level qualification label. If the room hears revenue, the finding has outrun its evidence whatever the words were.

The one thing you can honestly do with the gap is measure it separately, later. Stage attainment and eventual outcome are two different analyses, and a cohort that is too thin for the second today will not always be.

Starting at the MQL, layer one: the label is applied by your own team

The MQL step is larger than it looks, because two separate layers of your own operation sit between the buyer and the thing you are counting. Both are legitimate parts of how a business runs. Neither is a property of the buyer. Here is the first.

HubSpot defines a Marketing Qualified Lead as "a contact or company that your marketing team has qualified as ready for the sales team." Documented The definition names who does the qualifying and stops there. Across two readings of that documentation, no qualification criteria appear anywhere on it - not as a default, not as a suggestion. The criteria are yours.

The stage is also set by several different routes - manually on a record, in bulk from an index page, by an imported column, by a workflow or chatflow action, by a rule tied to a record's associations, and by integrations that sync the property. A field with that many entry points can carry values that were set for operational reasons rather than analytical ones.

The consequence for your analysis is direct: a finding about what precedes an MQL is partly a finding about your own scoring rules. If your model awards points for pricing-page visits, then pricing-page visits will be over-represented among MQLs, and discovering this is not a discovery about buyers. It is your scoring model reading itself back to you. That is a mechanical consequence of how the field is documented to work rather than something a vendor has measured, and the way to break the circularity is to check it directly rather than to reason about it. Recommendation Before running an MQL-based trace, write out your scoring or qualification rules and mark every exposure you intend to test that already appears in them. Those are not candidate findings. Test them separately, if at all, and say why they are excluded.

And one thing we cannot tell you. A deliberate search for research quantifying how far MQL definitions diverge between companies returned nothing - we found no study measuring it, in either direction. So the honest statement is narrow: the label is documented by vendors as a field the customer's own team populates, and we found no study quantifying how far definitions diverge between companies. That is a gap in the evidence, not a finding that definitions are uniform, and not a licence to assert that they vary wildly either.

Layer two: whether an MQL goes anywhere depends on your sales process

The second layer sits between the MQL and everything downstream of it, and there is published research on it.

Sabnis, Chatterjee, Grewal and Lilien studied what determines whether sales representatives follow up marketing-generated leads (Journal of Marketing, 2013). Their data comes from 461 sales representatives employed by four firms, and the outcome they measured is the proportion of a rep's time devoted to marketing leads rather than to self-generated ones. What they report is that this proportion "depends on organizational lead prequalification and managerial tracking processes… as well as marketing lead volume… and sales rep experience and performance," with experience and performance moderating the rest: as experience increases, responses to managerial tracking and to lead volume decrease while response to prequalification quality increases.

Read that as a statement about measurement rather than about sales management, and be careful about how far it reaches. The study suggests that sales attention to marketing-generated leads can vary with organisational processes, lead volume and representative characteristics. It did not measure whether an individual MQL was worked, converted or progressed - its unit is the representative and its outcome is how that representative divided their time. So applying those factors to your own MQL cohort is a hypothesis to test, not a result to import.

It is a hypothesis worth taking seriously all the same, because if attention to marketing-generated leads varies with process and with who received the record, then something sitting between the buyer's behaviour and your downstream count is varying too - and it is not a property of the buyer. Recommendation Test it on your own data rather than assuming it: compare downstream progression across teams, across periods and across reps of differing tenure, and see whether the MQL label behaves the same way in each. If it does not, that is a finding about your process, and it belongs in the write-up before any finding about buyers.

What an MQL cohort supports, then, is a real but narrow thing: a description of what preceded records your marketing team labelled as qualified, useful for auditing acquisition patterns and for finding where your own process is inconsistent. Recommendation Treat every MQL-based finding as a hypothesis about acquisition rather than a conclusion about buyers, and state the two layers in the write-up - the scoring rules that produced the label, and the fact that downstream progression depends on process factors you have not measured. A reader who knows both can still act on the finding. A reader who knows neither will act on it as though it were about buyers.

Lead level: an acquisition outcome, and nothing after it

At lead level you have observed acquisition - record creation is itself an outcome, and a real one - but you have not observed any downstream qualification, opportunity or revenue outcome. That distinction is the same endpoint discipline the rest of this page asks for, and it applies here too: if the question really is about acquisition, lead creation is a defined event with its own eligible cohort and its own comparison group. What the level cannot do is speak to anything after it.

What it can honestly do is describe: where volume came from, how the mix shifted, which sources are represented, and where the data quality is poor enough to distort everything above it. Those are useful things, and diagnosing them is often the fastest route to making the higher rungs usable later.

What it cannot do is distinguish the volume that becomes revenue from the volume that does not, because it never observes either. Recommendation When you present lead-level work, present it as a data and coverage audit and give it a heading that says so. The failure mode here is not a wrong conclusion; it is a correct description of volume that a reader converts into a conclusion about value on the walk back to their desk.

When the honest answer is "not yet"

Sometimes every rung fails: too few wins for the revenue question, an SQL definition that changed nine months ago, an MQL score that is mostly a proxy for the marketing you already do. "We cannot answer this yet" is a legitimate finding, and it is more useful than a confident answer built on a rung that cannot carry it - provided it comes with the two things that make it actionable.

The first is what specifically is missing. Not "the data is bad" but the particular shortfall: which comparison, at which rung, needs how much more of what before it becomes answerable. That turns an apology into a plan. The sufficiency arithmetic is what produces the number.

The second is what to start recording now. Fields that were never captured consistently cannot be reconstructed backward, but they can be captured from today - which is the part of the wait that does productive work. Recommendation Document the endpoint, the eligibility rule and the effective date before you analyse anything, and freeze that version within the cohort - not forever. Definitions legitimately change as segmentation, routing, products and sales capacity change; what breaks an analysis is an undated change, not a change. When the business revises a definition, create a new version, and then either analyse the periods separately or write down the mapping between them. Never apply today's definition silently to historical records. Alongside that: record the stage-entry timestamps you are currently overwriting; capture the exposure fields to the same standard on records that progress and records that do not; and set the cohort window for the analysis you intend to run in twelve months. The next analysis is only as good as the recording that precedes it.

A short cohort report in the meantime is not a defeat either: outcome states, denominators, unresolved shares and the traceable proportion, with no claims attached. It tells the board what is knowable, which is a different and more honest thing than telling them what you wish were known.

The sentence each rung lets you write

Everything above resolves into this, because the sentence is what actually leaves the room.

Outcome measured, and the strongest sentence it supports
OutcomeThe sentence you can writeThe sentence you cannot
Closed-won, mature cohort, comparable controls"In this cohort, deals that were won differed from comparable deals that were lost on X, by this much, with this interval"Anything causal, and anything about deals still unresolved at the cutoff
Reached the defined SQL stage"Among eligible records in this cohort, reaching the defined SQL stage was associated with X, by this amount and with this uncertainty" - a statement about qualification-stage attainment under the definition and capture rules in forceAnything about revenue, win rates or deal value. The measurement stopped before the outcome that would settle it
Reached MQL"Records our marketing team labelled qualified differed from comparable records that were not, on X, under the scoring rules in force during this window"That the finding is about buyers rather than partly about your scoring rules; that the label means the same thing across periods or companies
Lead level only"Volume, mix and source coverage looked like this, and here is where the recording is inconsistent"Anything separating volume that converts from volume that does not
No rung is sufficient"Under the stated baseline rates, allocation, test, alpha and power assumptions, the estimated target is N resolved records at rung R to detect a difference of D; here is what we are recording from today"A finding, of any kind, presented with the caveats in a footnote

None of the rows above is a failure. Each is a real claim that some analysis is entitled to make, and a team that consistently writes the row it earned will be trusted with the next question in a way that a team writing one row up will not.

And whichever row you land on, the finding is not yet a finding: a difference that survives the trace still has to survive selection, survivorship, multiplicity and halo before it means anything. That is the page on patterns and artefacts, and it applies to every rung on this ladder rather than only to the first.

Not sure which rung your cohort can actually carry? The question that settles it is not "how much data do we have" but "what difference would change a decision, and at which outcome." Everything else follows from those two answers, and settling them is a conversation rather than a project.

Book a Session


Sources

Sources. Lifecycle stage definitions: HubSpot Knowledge Base, "Use contact and company lifecycle stages," showing a last-updated date of 17 July 2026 and re-read on 4 September 2026 - a dynamic vendor page that may change without notice. A Marketing Qualified Lead is defined as "a contact or company that your marketing team has qualified as ready for the sales team," a Sales Qualified Lead as "a contact or company that your sales team has qualified as a potential customer," and an Opportunity as "a contact or company that is associated with a deal." Two readings confirmed that the page specifies no qualification criteria for either qualified stage. This is vendor documentation of one vendor's product behaviour, which the vendor controls and may change; it is not evidence about buyers, and the inference drawn here - that a finding about what precedes an MQL is partly a finding about the scoring rules that produced the label - follows from how the field is documented to work and is not something HubSpot has measured. Sales follow-up of marketing leads: Gaurav Sabnis, Sharmila C. Chatterjee, Rajdeep Grewal and Gary L. Lilien, "The Sales Lead Black Hole: On Sales Reps' Follow-Up of Marketing Leads," Journal of Marketing 77(1), 2013, pp. 52-67 - peer-reviewed, with data from 461 sales representatives at four firms. The measured outcome is the proportion of reps' time devoted to marketing leads rather than self-generated leads, modelled against organisational prequalification and tracking processes, marketing lead volume, and rep experience and performance. The widely repeated "70% of marketing leads are never followed up" figure appears in this paper as the framing clause defining the phenomenon it investigates, not as a result this study measured, and it is deliberately not quoted anywhere on this page. The study did not measure conversion, and reading it across to what an MQL cohort can support is an inference from a shared mechanism, labelled as such where it appears. Scoped absence: a deliberate search for research quantifying how far MQL definitions diverge between companies returned no study, in either direction. That is a gap in the published evidence and is stated on the page as such - it is not a finding that definitions are consistent, and not a basis for asserting that they vary widely. Censoring: T. G. Clark, M. J. Bradburn, S. B. Love and D. G. Altman, "Survival Analysis Part I: Basic Concepts and First Analyses," British Journal of Cancer 89(2), 2003, pp. 232-238 - a peer-reviewed tutorial in a clinical context, cited here only for the definition of right censoring as incomplete observation of a time to event. It is used on this page to mark a boundary rather than to license a method: censoring is a time-to-event concept, so it does not describe the unresolved records of a fixed-horizon binary comparison, which need a completed observation window instead. The application to CRM stage endpoints is CoreAEX's. Sample size: Geoffrey T. Fosgate, "Practical Sample Size Calculations for Surveillance and Diagnostic Investigations," Journal of Veterinary Diagnostic Investigation 21(1), 2009, pp. 3-14 - a peer-reviewed methods paper documenting the standard formula for comparing two proportions; its worked example concerns veterinary diagnostics, and it is cited here for one point only, that a required sample size is specific to a stated design and that reaching a calculated N does not guarantee a distinguishable result. The definitions of MQL and SQL used across this site, the rule that win rate is calculated only among closed outcomes, and the conversion-lag reasoning behind cohort timing belong to measuring content-market fit. The cohort-maturity rules and the sample-size arithmetic itself are set out on where to start a backward analysis, and the comparison-group construction on the backward-trace procedure; neither is re-derived on this page.

About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.

This page is part of the reverse funnel marketing cluster. The rest of the sub-pillars are in production and will be linked here as they publish.