Google and OpenAI each document that one prompt can become several queries. Neither documents how the results of those queries become an ordered list of vendors. That gap is why this page is about the buying question rather than about the machinery: the decomposition is something you can do, check and act on, and the selection step is not.
The practical claim here is modest and worth stating up front. Breaking a buying question into the sub-questions a buyer actually needs answered, then finding out which of them your evidence does not address, is useful work whether or not any engine ever runs those sub-queries. It surfaces gaps a prospect would hit anyway. This page is the decomposition part of the vendor-shortlist pillar, which owns what providers document about multi-query retrieval and what remains unpublished.
Two labels mark evidence boundaries on this page. Documented means the statement it sits beside is directly supported by the linked platform documentation, quoted. Recommendation means a method that CoreAEX prescribes and no platform documents. Research findings are attributed in the prose with their sample, dates, engine and the outcome they measured, and are not tagged. Untagged text is ordinary explanation.
What Providers Document, and Where It Stops
Two providers describe multi-query behaviour, in different vocabularies, on different products. The pillar sets both out in full; the short version matters here because it defines what a decomposition exercise can honestly claim.
Google's guide to optimizing for generative AI features defines the term directly: Query fan-out: A set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results to address the user's query.
Documented The same documentation describes grounding as relying on our core Search ranking systems
, and Google's AI-features page states that eligibility as a supporting link rests on ordinary indexing and snippet eligibility, with no additional technical requirements
.
OpenAI's help documentation describes ChatGPT rewriting queries and running follow-up searches, and states that placement is not guaranteed
. Documented
What none of them documents is the step this cluster cares about. How many sub-queries run, how their results are weighted against each other, how evidence from several of them is combined, or how any of that becomes a ranked list of products - none of that is published by any provider we located. And the observability boundary is specific rather than absolute. Google does not expose a complete list of the fan-out queries used for an AI Overview or AI Mode response, and neither provider promises a complete observable query trace - OpenAI documents that ChatGPT rewrites a question into one or more queries and may send further ones after reviewing results, without documenting or guaranteeing a complete user-visible query trace. So any externally generated or tool-captured set of sub-queries should be treated as partial or proxy evidence unless its collection method is documented. That constraint shapes everything below.
The Vendor Research That Speaks to This, and Its Two Limits
One vendor study we located asks the question directly - whether ranking for related sub-queries is associated with being cited - and it is worth reading closely, because it both supports and limits the same tactic.
Surfer SEO, an SEO software vendor, took the top-ten ranking pages for 10,000 keywords, found that 76% of those keywords triggered AI Overviews, and analysed 173,902 URLs. It reports that 51.2% of the AI Overview citations in its organic-ranking subset ranked for both the main query and at least one Gemini-generated related query, while 19.6% ranked only for the main query.
Those are composition shares, and the study's own headline converts them into something they cannot support. It states you're 161% more likely to get cited if you also rank for fanouts
- but 51.2% and 19.6% describe how the cited results were made up, not how often a page in each group got cited. Turning them into a page-level likelihood needs the number of eligible pages in each group as denominators, and those are not published. The composition finding stands; the 161% lift does not follow from it.
The article also reports the Spearman correlation here is 0.77
between its fan-out-ranking measure and its citation pattern. It does not publish the observation unit, the underlying series, any uncertainty or significance, or a reproducible methods section, so there is not enough on the page to reconstruct or interpret that coefficient. Treat it as a vendor-reported association rather than an effect estimate.
The first limit is where the fan-out queries came from. Google does not expose the queries its systems generate. The study says plainly: We pulled the top 10-ranking pages for 10,000 keywords. 76% had AI Overviews, so we used Gemini to pull their fan-out queries (33,000 in total).
So the related queries in this study were produced by a different model as a proxy, not observed from Google's AI Overview system. The finding is that pages ranking for a set of model-generated related queries were more often cited - which is genuinely informative, and is not the same claim as ranking for Google's actual fan-outs.
The second limit comes from a separate Surfer analysis, and it constrains the tactic more than the first study supports it. That second piece of work, based on 1,600 runs across a group of keywords and prompts, compared the generated terms across repeated runs - each run against the first, and each against the preceding one - and reports that only about 27% of fan out keywords actually stay consistent across different searches
. Across ten iterations it found that only a tiny 0.6% of keywords consistently appeared across all runs
, while 66% of fan-out keywords appeared only once
.
That analysis states no number of seed prompts, no geography and no collection dates. It also does not set out one collection procedure for all 1,600 runs: it describes fan-out queries returned by Gemini when grounding is used, and separately shows queries issued by ChatGPT, so the terms compared are not all from a single system. Read it as a scoped vendor observation rather than a general stability rate. Read at that scope, it still matters here: if most generated sub-queries differ between runs, a content plan built to target a specific fan-out list is aiming at a set that largely re-forms each time. The variation page covers how much else moves between runs.
The authors are direct about the design: correlation ≠ causation
, and our data only shows patterns, not that ranking for fanouts always improves your chances of getting cited in AIOs
. They also note we didn't pull SERP data beyond the top 10
, so the comparison runs among pages that already rank well. The page carries no collection dates, no geography and no separate methodology section, and covers Google AI Overviews only.
Read together, the honest summary is this. Among already-ranking pages, breadth of coverage across related queries was associated with citation, on one engine, in one vendor's dataset, using model-generated proxies for queries Google does not expose. That is a reason to care whether your evidence covers the parts of a buying question. It is not a reason to build a page per predicted sub-query.
Decomposing a Buying Question
Start from a question a real buyer would ask, in their words, and split it by the decisions inside it rather than by keyword variants. The method below is a CoreAEX procedure, not a documented one, and its justification is that it finds gaps in your evidence. Recommendation
Take a question of the shape a buyer actually types or says: "what's the best tool for managing customer onboarding for a 200-person B2B software company already using Salesforce?" Four distinct decisions are bundled inside it, and each one demands different evidence.
| Part of the question | What the buyer is deciding | Evidence it demands |
|---|---|---|
| "tool for managing customer onboarding" | Which category of product solves this at all | A clear statement of the category you are in, in the words buyers use for it |
| "for a 200-person B2B software company" | Whether you fit the size and shape of the buyer | Specific segment fit - team sizes, deployment scale, named customer profiles |
| "already using Salesforce" | Whether you work with what they have | Named integration, what it does, what it does not do, and how it is set up |
| "the best" | How you compare with the alternatives | Comparative evidence, much of which will not be on your own site |
Three rules keep this useful rather than mechanical. Recommendation
Split by decision, not by phrasing. "Onboarding software" and "customer onboarding tools" are the same decision in two vocabularies; "works with Salesforce" and "handles 200 seats" are two different decisions. Keyword tools cluster by phrasing, which is why a keyword cluster is a coverage input rather than a decomposition.
Keep the constraints attached. The size, the stack, the industry and the region are what turn a generic category question into this buyer's question, and they are often the parts your evidence does not reach. A page answering "what is customer onboarding software" answers none of them.
Stop at the decisions, not at exhaustiveness. A buying question typically contains three to six real decisions. Generating forty question variants produces a content backlog, not a decomposition.
Mapping Each Part to Evidence
The output of a decomposition is not a content plan. It is a map of what you can currently support, and the useful column is the one that comes back empty.
For each part of the question, sort your evidence into three categories. Recommendation
- Answered, on your properties. The fact exists, it is specific, and someone reading only your public pages could use it. Test that literally: try to answer the buyer's question using nothing but your own site and documentation, and note where you run out of specifics.
- Answered, but elsewhere. The evidence exists in a review profile, an independent comparison, a community thread or an analyst listing. The source-composition page covers which classes of source get cited and what each figure counts; in one small brand-level benchmark, external sources made up the large majority of the cited-source mix for SaaS brands.
- Not answered anywhere. This is the column worth having. A constraint no source addresses is a gap a buyer hits too, and it is actionable regardless of what any engine does.
One category depends on properties you do not own, and another is an evidence absence - which changes what the plan looks like. Some gaps need your own page or a documentation update; some need third-party coverage corrected or earned; and some first need a decision about whether you have a defensible claim at all. Comparative evidence in particular tends to live in third-party lists rather than on vendor sites, and the comparison-content page sets out what is and is not measured about those. Where a review profile carries the wrong tier or a stale integration list, the review-platform page and the consistency page cover the handling.
One research finding is worth attaching to this exercise, carefully. In a controlled study of how models select between conflicting evidence documents, Wan, Wallace and Klein found that current models rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone
. That study is about evidence selection on constructed conflicting documents, not about vendor recommendations. Read narrowly, it is a reason to check whether your pages actually address the question asked, rather than whether they look authoritative.
Why This Is Not a Page Per Query
The available mistake here is to treat a decomposition as a publishing schedule. It is worth naming, because the ingredients for it are all present: fan-out is documented, sub-queries can be generated in bulk, and one study associates breadth of ranking with citation.
Three things stand against it.
Externally generated queries are not necessarily the queries that ran. Google does not expose a complete AI Overview or AI Mode fan-out trace. A list generated separately by another model - as in Surfer's citation study - is therefore a proxy. Tool-captured queries should be described according to the tool's documented collection method and should not be assumed complete.
The list is unstable, on the same vendor's separate measurement. In the 1,600-run analysis described above - a separate piece of work from the citation study, published by the same vendor - about 27% of the generated terms recurred across repeated searches for the same prompt. A page built for a sub-query that appears in a minority of runs is a page built for a moving target.
Google warns against the tactic by name. Its generative-AI guide states: While it might be tempting to create separate content for every possible variation of how people might search (for example, by focusing on other queries that people have asked, or fan-out queries), doing so primarily to manipulate rankings or generative AI responses in Google Search violates Google's scaled content abuse spam policy.
It adds that this is also an ineffective long-term strategy, as a high quantity of pages doesn't make a website higher quality or more relevant to users
. Documented Note the condition: the policy turns on doing it primarily to manipulate rankings or responses. That is a policy reason to organise content around durable buyer decisions rather than predicted query strings.
And splitting one decision across six pages answers it six times partially, leaving a buyer to read all six.
The alternative is duller and holds up better: make sure the evidence exists, is specific, and sits where the question is actually being answered - which is often one thorough page, sometimes a documentation update, and frequently a third-party profile you do not own. Recommendation
What Is Not Known
No provider documentation located promises a complete query trace for a given response. Google does not expose a complete AI Overview or AI Mode fan-out list, while OpenAI documents query rewriting and additional searches without documenting that every resulting query will be visible. External parties therefore cannot reliably reconstruct the complete set of sub-questions used for a response.
Nothing located establishes that covering more of a buying question causes selection. The one study we located that addresses the neighbourhood is correlational by its authors' own statement, restricted to pages already ranking in the top ten, run on one engine, and built on model-generated proxies for the queries in question. It has no stated collection dates or geography.
The decomposition method on this page is not a validated protocol. It is a reasoned procedure whose justification is that it surfaces evidence gaps a buyer would encounter anyway. No study located tests decomposition as an intervention and measures anything downstream.
And the selection step remains unpublished. Google documents fan-out and grounding in its core ranking systems; OpenAI documents query rewriting and states that placement is not guaranteed. How retrieved evidence becomes an ordered list of vendors is not documented by any provider located, which is the boundary the whole cluster runs into.
Want to know which parts of your buyers' questions your evidence does not answer?
It is usually four or five decisions and one empty column - and that is a short conversation. Book a Session.
Sources
Sources: Platform documentation, quoted as read on August 31, 2026: Google Search Central, Google's Guide to Optimizing for Generative AI Features on Google Search (last updated 2026-07-10 - the query fan-out definition, the grounding statement, and the scaled content abuse passage quoted in full above, which is conditioned on publishing "primarily to manipulate rankings or generative AI responses") and AI features and your website (last updated 2025-12-10 - eligibility rests on ordinary indexing and snippet eligibility); OpenAI Help Center, Searching the web with ChatGPT - the article displays a relative update date rather than a fixed one, so it is cited to the review date. Neither provider documents how the results of multiple queries are combined into an ordered list of vendors. Vendor research, named as such, and two separate pieces of work rather than one. First, Surfer SEO, Joshua Hardwick, "Ranking for Multiple Fan-Out Queries Dramatically Increases Your Chances of Getting Cited in AIOs" (last updated December 6, 2025; top-ten ranking pages for 10,000 keywords, of which 76% triggered AI Overviews, and 173,902 URLs; the 33,000 fan-out queries were produced with Gemini rather than observed from Google, which the study states; Google AI Overviews only; no collection dates, no geography and no separate methodology section stated; the authors state "correlation ≠ causation", that the data "only shows patterns", and that "we didn't pull SERP data beyond the top 10"). Its 51.2% and 19.6% figures are composition shares of cited results; its headline "161% more likely" is quoted on this page only in order to reject it, because converting those shares into a per-page likelihood requires group denominators the study does not publish. Its reported Spearman coefficient of 0.77 is carried as a vendor-reported association, not an effect estimate, because the page publishes no observation unit, series, uncertainty or methods section. Second, and separately, Surfer SEO, Jakub Sadowski, "AI Search Study: Understanding Keyword Query Fan-out" (this page's own "last updated" field is unstable and returned different values across checks on the review date, so it is cited to the review date, August 31, 2026, rather than to any single displayed value) - the source of the 27%, 0.6% and 66% stability figures, from 1,600 runs across "a diverse group of keywords and prompts", comparing each run's fan-outs to the first run and to the previous run across ten runs; it states no seed-prompt count, no geography and no collection dates, and does not set out one collection procedure covering all runs, describing fan-out queries from both Gemini (with grounding) and ChatGPT. Peer-reviewed research: Alexander Wan, Eric Wallace and Dan Klein, "What Evidence Do Language Models Find Convincing?" (ACL 2024, Volume 1: Long Papers, pages 7468-7484; the ConflictingQA dataset pairs controversial questions with real-world evidence documents - a controlled study of evidence selection, not of vendor recommendations, and no figure from it is used on this page). No measurement cited here establishes that covering more parts of a buying question causes a vendor to be retrieved, cited, named or recommended.
About the author
Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.