An absence from one AI answer is a symptom, and a single observation almost never identifies its cause. The pillar sets out nine classes of explanation and what would distinguish each. This page is about the part that comes before and after that table: how to establish that you are genuinely absent rather than looking at noise, what order to run the tests in, what a server log actually licenses you to conclude, and what to do when the tests are exhausted and the omission survives.

One finding shapes the whole exercise. A system may identify a product when it is named and still omit it from an unbranded discovery answer. In Sharma's narrow API study, on a sample of recently-launched products, named-prompt recognition exceeded 94% while unprompted discovery remained below 9%. The study did not determine whether that recognition came from model memory, live retrieval or both.

That reframes the usual question. The operational one is whether the product appears in an unbranded discovery response - and, if it does not, which observable conditions can be ruled out. The omission may arise before retrieval, during retrieval, during response composition, or through provider-side selection, and no located study distinguishes those stages. That is why this page is a sequence of eliminations rather than an explanation.

Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked platform documentation, quoted. Recommendation means a procedure that CoreAEX prescribes and no platform documents. Research findings are attributed in the prose with their sample, dates, engine and the outcome they measured, and are not tagged. Untagged text is ordinary explanation.

First Establish That You Are Actually Absent

A "we are missing from AI answers" report often arrives as one screenshot, and a screenshot is a single draw from a distribution. Four checks come before any diagnosis, and each is cheap.

Check that you asked an unbranded question. Typing your product name and getting a competent description tells you almost nothing about whether you would appear when a buyer describes their problem instead. The two are different measurements, and the gap between them is the subject of the next section.

Check how many runs you have. One is not enough, and the published measurements of run-to-run variation are large enough that a single absence is close to uninformative. The variation page covers how much movement is ordinary; the measurement page covers the panel that makes a before-and-after answerable. In short: repeated runs of a fixed prompt set, per engine, per day, with conditions recorded.

Check whether an answer was generated at all. Google states that AI Mode doesn't always get it right and that for low-confidence responses it may return a set of web links instead of a generated answer. Documented A run that produced no generated answer is not evidence of your absence from one; it is a run that did not test the question, and it should be recorded as its own status rather than counted as a miss.

Check what you were measuring. Being absent from an answer, being present but unlinked, and being present and cited are three different outcomes, and teams routinely merge them. The pillar separates ten. If your report is "we are not in the shortlist," establish which of those outcomes it describes before anyone starts explaining it.

Only after those four is an absence worth diagnosing. A team that skips them will spend a quarter explaining variance. Recommendation

Being Known Is Not the Same as Being Found

Recognition and discovery are different outcomes, and one preprint measures both on the same sample. Separating them is worth doing before any diagnosis, because a team can hold strong evidence about one and none at all about the other.

Amit Prakash Sharma's preprint, submitted January 1, 2026, sampled 112 products from the top 500 featured on the 2025 Product Hunt leaderboard and ran 2,240 queries via API against two model endpoints: gpt-4o-mini and sonar with web search (Perplexity's API), between December 15 and 20, 2025. When products were named in the prompt, they were recognised in 99.4% of gpt-4o-mini queries and 94.3% of sonar queries. When the same products were sought through discovery-style questions - the paper's example is "what are the best AI tools launched this year?" - the rates were 3.32% (gpt-4o-mini) and 8.29% (sonar). The author describes the gpt-4o-mini gap specifically as a gap of 30-to-1; the sonar gap works out to roughly 11-to-1.

The scope on that finding is narrow and has to travel with it. Both endpoints were queried via API - model endpoints, not the consumer ChatGPT and Perplexity products a buyer would use. The sample is newly-launched Product Hunt products, which the author notes skews toward certain product categories (developer tools, productivity, AI) and may not describe consumer products, enterprise software, or products launched through other channels. Two models were tested; the author writes that Claude, Gemini and others might show different patterns.

Even held to that scope, the direction is the useful part. Recognition and discovery are separate outcomes, and nothing in the study connects movement in one to movement in the other. A team whose evidence is that the system "knows who we are" has tested the outcome that was never in doubt.

The same paper reports a null that deserves care. Its composite GEO score - a regex-based measure of on-page optimisation for AI visibility - showed no statistically significant correlation with discovery on either endpoint: r = −0.108, p = 0.256 for gpt-4o-mini and r = −0.102, p = 0.286 for sonar. That is a study reporting no statistically meaningful relationship detected, which is a different and stronger statement than no evidence existing. It is also one study, on one sample, with a scoring method the author names as a limitation: GEO scoring was regex-based. More sophisticated content analysis might capture signals my approach missed. It does not establish that on-page work is useless; it establishes that this measure of it showed no statistically significant association with discovery in this sample.

The result that most complicates this page is what the paper found when it looked for what was associated with discovery. On the sonar (Perplexity) endpoint, the paper reports seven statistically significant correlations in total, including referring domains at r = +0.319 (p < 0.001), Product Hunt ranking at r = −0.286 (p = 0.002), and Reddit mentions (cleaned) at r = +0.395 (p = 0.002). The "cleaned" Reddit variables exclude the 52 of 112 products (46%) whose names were too generic to search reliably - a removal the author notes reduces statistical power for the community signal analysis. These are bivariate associations within this sample. They are not demonstrated causal effects, and nothing here establishes that they would predict discovery out of sample.

On the gpt-4o-mini (ChatGPT) endpoint, none of the variables tested showed a statistically significant correlation with discovery. The author characterises that as discovery appearing essentially random. More narrowly, the study detected no statistically significant association among the variables it measured in this sample - which is not the same as randomness, and not the same as showing that no association exists. Unmeasured variables, limited statistical power and the study's design all remain available explanations.

That is a genuine tension with the premise of a diagnostic procedure, and it should be stated rather than resolved. On one endpoint, in one week, on 112 recently-launched products, none of the variables tested showed a statistically significant correlation with which products surfaced. The honest reading is not that diagnosis is pointless - the study measured a specific set of on-page and off-page variables on a specific sample, not the full set of reasons a mature B2B vendor might be absent. But it is a reason to hold any single explanation loosely, and a reason the last section of this page exists.

Ordering the Diagnosis by What Each Test Rules Out

The pillar's table lists the classes of explanation; the useful addition is the order. Run the tests that are cheap and that eliminate whole classes first, so that expensive investigation is spent only on what survives.

Three tiers, in this order.

First, the documented necessary condition. Google states that to appear as a supporting link in its AI features a page must be indexed and eligible to appear in search snippets, and that there are no additional technical requirements. Documented That single sentence does two things: it gives you one condition you can verify directly, and it tells you that no further technical prerequisite is documented for that surface. Confirm indexation and snippet eligibility first, because failing it explains everything downstream and passing it removes an entire class of theory. Note the boundary: this is Google's documentation about Google's features, and no equivalent statement was located for other providers.

Second, the checks you can run entirely on your own material. These need no engine and no panel. Search your own site for the category term the prompt uses. Try to answer the prompt yourself using only your public pages, and see where you run out of specifics. Compare the same handful of product facts across your site, your documentation and your third-party profiles. Each of these is an hour's work and each either eliminates a class or produces a concrete defect to fix. They are also worth doing on their own merits - a buyer hitting the same missing specifics has the same problem - which is why they are worth doing before anything has been proven about their effect on retrieval. Recommendation

Third, the checks that need a panel. For comparisons across prompts, regions, accounts or engines - with and without a constraint, one location against another, one engine against another - use repeated matched runs under recorded conditions. Run-to-run variation can otherwise obscure the difference being investigated, or imitate one that is not there. Report each condition separately rather than as a difference, and specify the comparison before inspecting the results. This is the expensive tier, and the point of the first two is to arrive here with fewer hypotheses.

Two documented facts sit across the whole sequence and change what a comparison can mean. OpenAI's help documentation states that ChatGPT may use saved memories when rewriting a search query, and that IP-derived and saved location may influence results. Documented So a colleague in another country, or logged into a different account, is not running your test. Record location and session state as conditions, or the comparison is not a comparison.

What a Log Line Does and Does Not Show

Server, CDN and edge logs are easy to over-read. They are genuinely useful, and what they establish is narrower than what teams routinely conclude from them.

A request recorded in a log shows one thing: a request reached the layer that did the logging. It does not show that the page was rendered, that the content was stored, that anything was retrieved for a later answer, that your page was used in composing one, that it was cited, or that your product was recommended. Each of those is a separate event, and none of them appears in a web server's access log.

The absence of a request is worth more than its presence, and even then only as elimination. If the pages you expect to matter show no verified requests from the relevant documented agent within the logs and window you actually retain, investigate access, attribution, cache and CDN coverage, and bot-specific delivery. That absence narrows the evidence available to you; it does not by itself show that the provider could not reach or retrieve the page by another route - a different documented agent, a search index, cached content, an unmonitored IP range, or traffic that never reached the layer you log. The investigation it triggers belongs to the AI-crawler guide, which owns access, rendering and bot-specific delivery. This page does not re-derive that work; it only marks the point at which a diagnosis hands off to it.

Three habits keep log evidence honest. Separate indexability, fetchability and bot-specific delivery as three findings, not one - a page can be indexable and undelivered, or fetched by one agent and blocked to another. Treat a user-agent string as a claim, not an identity, since it is trivially set by anything; the engines that publish verification methods are the ones whose traffic you can actually attribute. And state the window: "no fetches" over three days and over three months are different findings. Recommendation

When the Answer Is "Unexplained"

Working the full sequence and finding nothing is a legitimate outcome, and recording it as such is more useful than the alternative.

The pillar's last class of explanation - provider-side selection behaviour - cannot be distinguished from outside, because no provider publishes how retrieved evidence becomes an ordered list of vendors. When the testable classes are checked and the omission survives, the finding is that the cause is unknown. It is not that the cause is whichever theory the room finds most interesting.

Two things make an unexplained result worth having. It is a record that the cheap explanations were eliminated, which stops the same four hypotheses being re-proposed every quarter. And it is a baseline: an omission that persists across a full panel over several weeks is a different object from one screenshot, and it is the only version of the observation worth escalating.

What follows from an unexplained absence is ordinary work rather than a tactic. Publish the specifics a comparison would need. Keep your third-party listings accurate. Make sure your pages can be fetched and rendered. Each of those is justified by readers and buyers using them, and none of them is a measured lever on shortlist inclusion - which is exactly why they survive an engine update, and why they are the right thing to do while the cause is unknown. Recommendation

The Diagnostic Shortcut, and What It Costs

The costly failure here is usually not a wrong diagnosis; it is a fast one. The pattern is familiar: a vendor is missing from one answer, someone proposes a cause, and the proposed cause happens to match a purchase or a project the organisation was already inclined toward.

Three moves make it recognisable.

A single observation is treated as a measurement. One prompt, one session, one day, and a conclusion about standing. Against measured run-to-run variation, that observation is consistent with a real absence and equally consistent with an ordinary fluctuation.

An explanation is selected before the eliminations are run. The order in the previous section exists because it is cheap; skipping it means the surviving hypothesis is the one someone liked, not the one that survived.

A remedy is bought and then validated against the same weak observation. If one screenshot established the problem, another screenshot will establish the fix, and neither is evidence. A change is measured with the same panel design that measured the absence, or it is not measured.

None of this argues for inaction. It argues that the question "why are we missing" deserves the same evidential standard as any other claim the business acts on, and that the standard is not high - repeated runs, recorded conditions, and eliminations before explanations.

What Is Not Known

No provider publishes how retrieved evidence becomes an ordered list of vendors, so the final class of explanation on the pillar's table is permanently outside a customer's reach. Google documents that its AI features are grounded via its core Search ranking systems and that eligibility rests on ordinary indexing; OpenAI documents query rewriting and states that placement is not guaranteed. Neither describes the selection step, and no located documentation from any provider does.

The relationship between what a model retrieves and what it finds persuasive is only partly studied, and not on this question. Wan, Wallace and Klein found in a peer-reviewed ACL 2024 study that current models rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone. That is a controlled head-to-head over constructed conflicting evidence, not a study of vendor recommendations, and it should be read as a reason to check whether your pages actually address the question asked - not as a finding about shortlists.

Absence has not been studied as a subject. The evidence used on this page measures discovery rates, retrieval overlap, variation and documented eligibility. We located no study that takes a set of vendors known to be absent, works through candidate causes, and reports which ones accounted for the absence. The diagnostic order set out here is therefore a reasoned procedure built from what each source can eliminate, not a validated protocol.

And one study found no statistically significant association on one endpoint. On the gpt-4o-mini (ChatGPT) endpoint in Sharma's sample, none of the variables tested showed a statistically significant correlation with discovery. Two models, 112 recently-launched products, one week, via API. It is a narrow result and it may not describe an established B2B vendor at all - but it belongs beside a page recommending a diagnosis, because it is a published reason to expect some absences to remain unexplained after competent work.

Missing from answers and not sure where to start?

The first hour is usually elimination rather than investigation - and that is a short conversation. Book a Session.

Sources

Sources: Platform documentation, quoted as read on August 31, 2026: Google Search Central, AI features and your website (last updated 2025-12-10) and Google's Guide to Optimizing for Generative AI Features on Google Search (last updated 2026-07-10); Google Search Help, AI Mode in Google Search; OpenAI Help Center, Searching the web with ChatGPT - the OpenAI article displays a relative update date rather than a fixed one, so it is cited to the review date. Academic research: Alexander Wan, Eric Wallace and Dan Klein, "What Evidence Do Language Models Find Convincing?" (ACL 2024, Volume 1: Long Papers, pages 7468-7484; peer-reviewed; the ConflictingQA dataset pairs controversial questions with real-world evidence documents - a controlled study of evidence selection, not of vendor recommendations, and no figure from it is used on this page). Preprint, not peer-reviewed: Amit Prakash Sharma, "The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries" (arXiv 2601.00912v1, submitted January 1, 2026; 112 products sampled from the top 500 featured on the 2025 Product Hunt leaderboard, 2,240 queries executed via API against gpt-4o-mini and sonar with web search between December 15 and 20, 2025 - model endpoints rather than the consumer ChatGPT and Perplexity products; on the sonar endpoint the paper reports seven statistically significant correlations with discovery, versus zero on gpt-4o-mini; the "cleaned" Reddit variables exclude the 52 of 112 products, 46%, whose names were too generic to search reliably; the author's stated limitations are the Product Hunt sample skew toward developer tools, productivity and AI products, that only two models were tested, that excluding those products reduces statistical power for the community-signal analysis, and that GEO scoring was regex-based). Run-to-run variation figures referenced on this page are set out in full on the variation and measurement pages linked above, with their samples, intervals and limits. No measurement cited here identifies the cause of any vendor's absence from an AI answer, and no source located studies absence as a subject.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.