Two labels mark evidence boundaries on this page. Documented marks a platform definition or behaviour directly supported by the linked provider documentation, whether quoted or faithfully described. Any surrounding interpretation remains CoreAEX analysis unless separately attributed. Recommendation means a workflow, threshold or ownership rule that CoreAEX prescribes and no source specifies. Untagged text is ordinary explanation, or a conclusion following from something already labelled - and a conclusion drawn from documentation is not itself documented, which matters more on this page than on most.
A backward trace runs on fields. Lifecycle stage, lead source, campaign, attribution credit, opportunity stage, close date, amount. Before any of it means anything, you need to know what those fields actually hold - which is a different question from whether they are filled in correctly, and a harder one to answer from inside a report.
This page is the audit that comes before the analysis. It is not a data-cleaning exercise, and it contains no instructions for configuring anything. It is a set of questions you can answer about your own system in an afternoon, using your vendor's own documentation and your own records.
Two things to establish, and only one is "data quality"
The first is what the fields record. Vendors document this plainly, and what they document is narrow: these are labels your organisation configures and credit rules your organisation selects, not measurements of buyer behaviour.
The second is whether they were recorded to the same standard on the deals you won and the deals you lost. This is a different question from whether the fields are correct, and it is the one that decides whether a comparison between wins and losses means anything at all - every field can be individually accurate and the comparison can still be broken.
Both have to hold. The first without the second gives you well-understood fields and an invalid comparison; the second without the first gives you a symmetric comparison between two things you have misnamed.
A lifecycle stage is a label your team applies
Start with the field most often treated as though it described the buyer.
HubSpot's documentation defines a Marketing Qualified Lead as "a contact or company that your marketing team has qualified as ready for the sales team" and a Sales Qualified Lead as "a contact or company that your sales team has qualified as a potential customer," adding that this stage "includes sub-stages that are stored in the Lead Status property." A Customer is "a contact or company with at least one closed deal." Documented
Read those definitions for what they do not contain. There are no qualification criteria in them, because the criteria are yours. The vendor documents the existence of the label and who applies it. Everything the label means at your company was decided in a meeting.
The same page documents two separate sets of routes by which the value gets set, and they are worth keeping apart. The first is user-initiated: updating the property manually on a record, bulk-updating from an index page, importing a Lifecycle stage column, a chatflow action, a workflow action, or "certain integrated apps that sync the Lifecycle stage property, such as the HubSpot-Salesforce integration." The second is automatic, configured in settings: a default stage for newly created records, updates based on a record's associations, and a default stage for records synced from each connected app. Documented
Six ways to update a record, plus three categories of automatic setting. Those routes can all implement one agreed lifecycle definition, or they can encode conflicting criteria - the count itself proves neither. What it establishes is a provenance and governance risk: the more routes in use, the more places an inconsistent rule can enter without anyone noticing. Audit the rule applied by every route actually in use, rather than assuming the field means one thing because it has one name.
The practical conclusion - and this is an inference from how the field is documented to work, not something the vendor has measured - is that an MQL count should not be assumed comparable across two companies, or across two periods at the same company once the definition or the workflow behind it has changed. That is not a criticism of the field. It is a statement about what kind of object it is.
HubSpot's automated lifecycle updates normally move forward
This one has direct consequences for a backward trace, and it is easy to miss because the field presents as a simple status. It also needs stating precisely, because the constraint is narrower than it first appears. The HubSpot behaviour described throughout this section is quoted or faithfully described from its current documentation; the readings drawn from it are ours. Documented
HubSpot documents that "the default Lifecycle stage property can only be moved forward by HubSpot tools (e.g., import, form submission, API, Salesforce integration, and workflows). You must clear the value manually or via a workflow before using these tools to set an earlier value in the Lifecycle stage order." The association-based automation carries the same constraint: "these settings will not set lifecycle stage values backwards," illustrated with the case where "if a contact has a lifecycle stage of Customer, and a deal is created and associated with that contact, the lifecycle stage of the contact will not change to Opportunity."
When the stage is updated only through those documented tools and automatic settings, the current value normally behaves like a high-water mark - but it is not inherently a current-status field, and it is not inherently a one-way one either. A user can move it backward manually, or clear it before an automated reset, so the first thing to establish is which update routes your own portal permits and actually uses. Three consequences follow for anyone tracing backward, none of them documented by the vendor and all of them ordinary readings of the mechanism:
- Where only the automated routes are in play, the current value tells you the furthest a record ever got rather than where it is now. An account that reached Customer and churned still reads Customer.
- The current value alone does not reconstruct the path - but HubSpot documents properties that help. It records stage-specific calculated properties for "Date entered [stage]" ("the date and time when the contact or company entered the stage"), "Date exited [stage]", "Latest time in [stage]" and "Cumulative time in [stage]," noting that "a Professional or Enterprise subscription is required to use the latest time and cumulative time properties." Confirm which of these your reports actually use, and whether you also need the underlying property history for provenance.
- Manual backward movement does not behave uniformly across those properties, and the direction is easy to get wrong. HubSpot states that if you manually set a stage to an earlier value, "the legacy Became a [lifecycle stage] date property corresponding to the greater value will be cleared, and the property corresponding to the new lesser stage will be automatically populated" - while "the new calculated properties will not be cleared when you move a lifecycle stage backwards." Legacy and calculated properties therefore tell you different stories about the same record.
An attribution number is a credit rule someone chose
An attribution number combines recorded interactions with an attribution model and its configuration. The customer selects or configures the model; the allocation itself may be rule-based or computed from historical data. Either way the resulting credit answers a model-defined question - it is not a direct measurement of what caused the buyer to convert, and software having calculated it does not make it one. Everything the vendors state about their own products in this section is quoted or faithfully described from their documentation; the reading of it that follows is ours. Documented
Salesforce documents Campaign Influence as "a tool that helps you attribute a percentage of success to influential campaigns," which "identifies revenue share with standard and custom attribution models that you can update manually or via automated processes," and instructs the administrator to "choose a predefined model or enter custom influence percentages." The customer selects or configures the model; depending on which model, the credit may be entered by hand, calculated by a rule, or computed from historical data. Salesforce documents the last of these too: its Einstein Attribution model for Account Engagement "is based on a concept borrowed from cooperative game theory called the Shapley Value," and "uses your historical data to identify common journeys that users take," scanning campaigns that influenced an opportunity and assigning "a marginal contribution amount to each touchpoint" - available, per the documentation, in "Account Engagement Advanced and Premium Editions with Salesforce Enterprise, Performance, or Unlimited Edition." None of that makes the output a measurement of cause; it makes it a different kind of estimate.
Google Analytics 4 currently offers three attribution models: data-driven attribution, which "distributes credit for the key event based on data for each key event"; paid and organic last click, which "ignores direct traffic and attributes 100% of the key event value to the last channel that the customer clicked through (or engaged view through for YouTube) before converting"; and Google paid channels last click. The same documentation records that "the first click, linear, time decay, and position-based attribution models are no longer available as of November 2023." Note Google's term is key event, not conversion - a distinction worth preserving when you reconcile its numbers against anything else.
HubSpot documents nine models - linear, first interaction, last interaction, U-shaped, W-shaped, time decay, full path, J-shaped and inverse J-shaped - and describes their function this way: "attribution models attribute credit to the interactions that created contacts, deals, and revenue in HubSpot, and will apportion higher credit to key conversion points in the lead conversion journey."
One gap in that documentation is worth naming, because it is the term everything else rests on. The page describes what interactions do and lists types of them, but contains no definition of the form "an interaction is…". It documents eleven interaction types enabled by default for revenue attribution - among them page viewed, form submitted, CTA clicked, marketing email clicked, call connected, meeting attended, sales email reply, conversation, ad clicked, social post clicked, and contact created in HubSpot - with five more available but off by default, including media plays and marketing-event attendance. Both lists are presented as complete for this feature, not as a partial sample. What is missing is not a fuller enumeration; it is a definition of what qualifies as an interaction in the first place, on this page or any other we found.
HubSpot also documents two processing choices that change which interactions reach a report. Administrators can enable or disable interaction types, and "once you've made changes to your interaction types, your existing reports will need to be reprocessed. This may take up to two days, depending on the amount of data in your HubSpot account." And the model samples above a stated ceiling: "using sampling, HubSpot's attribution model can process up to 100,000 associations or activities per deal. If a deal is linked to more than 100,000 interactions - such as page views, calls, meetings, or other engagements - up to 100,000 of those interactions are used in the attribution analysis." Documented HubSpot presents this as a very-high-volume case, but does not publish how frequently deals reach the threshold. The practical step is to record which interaction types were enabled when an export was taken, and to check your own cohort for deals at that activity level, before treating two exports as comparable.
Two vendors, three models and nine, four models withdrawn from one of them in a single month, and a foundational term left undefined on the page that uses it most. None of this makes any of the numbers wrong. It makes them reporting constructs, answering model-defined questions, and it means the number you inherit depends on a choice somebody made from a menu.
Two systems, two different views of the same business
There is a second reason two attribution reports on one business disagree, and it is larger than the choice of model: the systems may not be looking at the same events, or at the same people. The three statements of GA4 behaviour below are from Google's current documentation; the consequence drawn between them is ours. Documented
GA4 documents three reporting identities. Blended works "by User-ID, device ID, then modeling - uses the user ID if it is collected. If User-ID information is not available, then Analytics uses the device ID. If no identifier is available, Analytics uses modeling." Observed is "by User-ID, then device ID," without the modelling fallback. Device based "uses only the device ID and ignores all other IDs that are collected." A CRM, by contrast, resolves identity through records and associations. These are different unit-of-analysis definitions, and neither is a version of the other.
GA4 also documents finite lookback windows. For acquisition key events - first_open and first_visit - "the default lookback window is 30 days," switchable to 7. "For all other key events, the default lookback window is 90 days," with 30 or 60 also available.
Hold that against a B2B sales cycle. Where a cycle runs longer than the configured window, touchpoints at the start of the buying process fall outside the window the model can see - so the report is not describing the same span of time your backward trace is. That is an inference from the documented defaults rather than a finding, but it is arithmetic rather than speculation, and it is checkable against your own median cycle length in an afternoon.
And attribution does not reach every dimension in the report. GA4 states that "user- and session-scoped traffic dimensions, such as Session source or First user medium, are unaffected by changes to the reporting attribution model" - which apply instead to event-scoped traffic dimensions. Two dimensions sitting beside each other in one report can therefore respond to a model change differently, and a reader comparing them without knowing that will conclude something false about both.
What "closed won" actually is in your configuration
Salesforce documents the OpportunityStage object as one that "represents the stage of an Opportunity in the sales pipeline, such as New Lead, Negotiating, Pending, Closed, and so on." Two boolean fields on that object carry the part a backward trace depends on. Both field descriptions below are quoted from that reference. Documented
IsClosed "indicates whether this opportunity stage value represents a closed opportunity (true) or not (false)" - and the description continues: "multiple opportunity stage values can represent a closed opportunity." IsWon "indicates whether this opportunity stage value represents a won opportunity (true) or not (false)," with the same addition: "multiple opportunity stage values can represent a won opportunity."
Those two trailing sentences are the operationally important ones. Won-ness is a flag carried by a stage, not a name - and the documentation states plainly that more than one stage can carry it. So a query that identifies won deals by a display label rather than by the flag will undercount wherever a second won stage exists, and a second won stage is exactly the kind of thing that gets added quietly for a renewal motion, a self-serve tier or a regional process.
The practical instruction, therefore, is not to trust the label. Recommendation Inspect or query the configured stage mapping and establish which stages in your own process have IsWon set to true. Do not identify won opportunities by one display name unless you have confirmed it is the only stage whose flag is set. That is a fact about your configuration, visible to you and invisible to us, and it takes a few minutes to settle.
A note on this citation. The IsClosed and IsWon field descriptions above were confirmed against Salesforce's Object Reference documentation at API version 190.0 (linked above), which renders the field content directly. A fully unversioned form of the same reference page returns navigation only, with no field content, through our tooling - so the versioned link is the one to check if you verify this yourself. No claim is made anywhere on this page about any other field on the OpportunityStage object.
The check that decides whether the comparison holds
Everything above concerns what a field means. This section concerns something else entirely, and it is the precondition a backward trace actually depends on.
Bernard Forgues, setting out the conditions under which selecting cases on the outcome yields valid inferences (Strategic Organization, 2012 - a methodological essay, not a study with its own data), states the requirement this way: "following the comparable accuracy principle, independent variables should be measured with the same accuracy in controls as they are in cases, unless the effect of inaccuracy is controlled in analyses." He also notes that the problem is frequent because cases tend to be better documented than controls, a tendency he links to retrospective data collection.
That describes most CRMs precisely. Deals that progress accumulate richer records: more logged calls, more notes, more campaign associations, more complete firmographics because someone did the research when the deal looked real. Deals that died in week two have a form fill and a disposition code. Compare the two directly and any dimension that correlates with effort spent will separate wins from losses beautifully, regardless of whether it had anything to do with the outcome.
Note the exception in Forgues's own sentence, because it changes what you have to do about this. The requirement is not that measurement must be equal - it is that unequal measurement must be either fixed or accounted for in the analysis. That is a materially easier bar, and it means a trace over asymmetric records is not automatically void. What it cannot be is silent: an analysis that compares a rich record against a sparse one without saying so is presenting an artefact of record-keeping as a finding about buyers.
The measurement policy that has to sit underneath any influence number - eligible audience, qualifying touch, lookback window, identity and account-stitching rules, opportunity-association rule, deduplication - is set out on measuring content-market fit, which owns it for this site. This page adds one requirement on top of that list: every one of those policy decisions has to be applied identically to the deals you lost - and the check that matters is per-field, not aggregate. Forgues's principle concerns the accuracy with which each independent variable is measured in cases and controls, so the audit below works field by field rather than counting fields. Recommendation
The audit, on your own data
Nine questions. None of them requires a tool you do not already have, and the answers are yours rather than ours. Recommendation
- For each stage in your lifecycle or pipeline, write down the criteria in one sentence. If two people write different sentences, that is the finding - and it is more common than the existence of a documented definition would suggest.
- List every route by which each stage gets set, manual and automatic, and note which routes are actually in use today rather than which exist.
- Establish whether your stage field moves backward, and if it does not, find out where the timestamped stage-change history lives and how long it is retained. That finding also feeds a decision upstream of this one: where to start a backward analysis depends in part on how much reliable stage history you actually have.
- Name the attribution model your report is using, and the date it was last changed. A model change mid-period makes a year-over-year comparison a comparison of two methods.
- Compare your median sales cycle against your analytics lookback window. If the cycle is longer, know which touchpoints are structurally outside the window.
- Read your own opportunity stage list and confirm how many stages can end a deal as won.
- Take a screening sample of closed-won and comparable closed-lost opportunities, and go field by field. For each field you actually intend to analyse, measure three things on both sides using the same observation window or stage cutoff: how often it is missing, where the value came from, and when it was recorded. A total count of populated fields can flag asymmetry worth investigating, but it does not establish comparable accuracy - some blanks are simply inapplicable, some populated values are wrong, and won deals mechanically accumulate more fields because they stay open longer and reach later stages. Account for deal duration, stage reached, and whether the value was recorded before the outcome was known.
- Identify which fields were written by a human who knew how the deal was going - qualification notes, engagement scores, sentiment, relationship strength - and separate them from fields a system stamped before the outcome was known.
- Write down, for each field you intend to analyse, whether it was recorded the same way on both sides. Where it was not, decide now whether you will account for the difference or drop the field.
In our work, the field that most often turns out to mean something different from what the reporting assumes is lifecycle stage - usually because the criteria were rewritten at some point and the historical records were never revisited. That is an observation from client engagements rather than a measured rate, and your own answer to question 1 will tell you more than ours would.
What is not known about CRM data quality
It would be convenient to end with a figure - some percentage of CRM records that are incomplete or wrong. There is no honest one to give.
We searched for independent, non-vendor research measuring the accuracy or completeness of attribution and touchpoint data in production CRM systems, and found none. The search covered peer-reviewed and independent studies of CRM attribution and touchpoint record accuracy; the origin of the widely circulated "CRM data decays at 30% per year" figure; academic literature on measurement error and validity in multi-touch attribution, including missing touchpoints; empirical field studies of enterprise CRM data quality in the information-systems literature; studies quantifying missing lead-source or campaign fields on opportunities; and independent benchmarks of attribution data completeness published between 2024 and 2026.
What that search returned is worth reporting precisely. Every CRM-decay percentage we were able to trace originated with a company selling data-quality or data-enrichment services, and none of those published a methodology, a sample definition, or an independent replication. The genuine academic literature on CRM data quality exists but answers a different question: it addresses data quality as a construct, as a governance problem, and as something users perceive - not whether the attribution fields in a live CRM are actually correct.
So this page states neither that CRM data is too poor for this analysis nor that it is good enough. Both would be claims nobody has evidence for. What can be said is narrower and more useful: the quality of your own records is measurable by you, question 7 above measures the part that matters most for a backward trace, and it takes an afternoon.
If that audit comes back badly - losses recorded far more thinly than wins, stage criteria nobody agrees on, an attribution model changed mid-period - the finding is about your instrumentation rather than your funnel, and it is a real result. It tells you what the next trace can and cannot be asked to answer, which is more than a confident analysis over unexamined fields would have done.
Not sure your fields say what your reports assume? Start with the fields you actually plan to analyse, and compare their missingness, provenance and recording time across won and comparable lost deals. The result tells you which variables can support a backward trace, which need adjustment, and which should be left out of it.
Sources
Sources. All vendor documentation was read on September 4, 2026, and all of it describes product behaviour the publisher controls and may change; none of it is evidence about buyer behaviour, and this page contains no configuration guidance for any product. HubSpot Knowledge Base, "Use contact and company lifecycle stages" (page displays last updated July 17, 2026) - the stage definitions, the six user-initiated routes, the three automatic routes, and the forward-only constraint; and "Automatically set and sync record lifecycle stages" (last updated July 13, 2026) - the association-based automation, whose backward-movement statement is scoped to those settings specifically. HubSpot Knowledge Base, "Understand attribution report definitions" (last updated June 21, 2026) - the nine models, the credit sentence, and the 100,000-associations sampling statement; and "Select interaction types for your attribution reports" (last updated December 31, 2025) - the interaction types enabled and disabled by default for revenue and deal-create attribution, and the reprocessing statement. The asset-type list on that page is deliberately neither quoted nor counted: two independent reads returned different groupings and a stated total that did not match the items enumerated. Google Analytics Help, "Get started with attribution", "Reporting identity", and "Select attribution settings" - none of these three displays an update date, so all are cited to the review date above. Salesforce Help, "Customizable Campaign Influence" - no update date displayed. Salesforce Developers, Object Reference, OpportunityStage (API version 190.0) - the object description and the IsClosed and IsWon field descriptions, confirmed verbatim on this versioned reference page; a fully unversioned form of the same reference returns navigation only through our tooling, so the versioned link above is the one to check. No claim is made anywhere on this page about any other field on that object. Salesforce Help, "How Einstein Attribution Works" - the Shapley-value basis and the edition requirement; note the documentation describes the mechanism in those terms and this page does not attribute the phrase "data-driven" to it. Research methodology: Bernard Forgues, "Sampling on the dependent variable is not always that bad: Quantitative case-control designs for strategic organization research," Strategic Organization 10(3), 2012, pp. 269-275 - a peer-reviewed methodological essay presenting no original data; the comparable-accuracy principle is quoted in full including its exception, which is the part that determines what a practitioner has to do about asymmetric records.
About the author
Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.
This page is part of the reverse funnel marketing cluster. The rest of the sub-pillars are in production and will be linked here as they publish.