Two labels mark evidence boundaries on this page. Documented means the guidance that follows is directly supported by published Schema.org or Google documentation, quoted. Recommendation means a modelling choice, threshold or workflow that CoreAEX prescribes and the documentation does not. Statements from platform staff in interviews, forums or podcasts are attributed in the prose rather than tagged - they are first-party, but they are not documentation. Untagged text is ordinary explanation or a conclusion following from something already labelled.

What We Found

We published four template-matched pricing pages representing four source-layer configurations, then asked five AI systems to read each one. Each configuration is tied to its own fictional vendor, feature copy and price value, so this is a controlled diagnostic comparison rather than a randomised or counterbalanced experiment - the layer treatment was not rotated across shared page identities. Two results, both from three engines running three fresh sessions each on August 30, 2026:

  • A price present only in JSON-LD, with no price anywhere in the rendered text, was reproduced in 0 of 9 runs. In nine separate access-check sessions covering the same engine-and-page combinations, all three engines described a distinctive feature of that page accurately. Those checks establish that the engines could reach and read the page during the observation window. They do not establish that each individual pricing response retrieved the page's content, because the check ran in a different session.
  • Where the visible text and the JSON-LD stated different prices, every number returned was the visible one. The markup value appeared in none of the nine runs, and no engine mentioned the discrepancy.

Two further interfaces - Claude and Perplexity - produced no usable page data and failed every access check, so they are excluded from both denominators. An access failure measures availability, not markup extraction.

The practical reading is narrow and it is the one we would give a client: if a commercial fact exists only in your structured data and nowhere a reader can see it, this test gives no reason to expect it to reach an answer. That is a statement about three named engines on one date under URL-supplied conditions. It is not a statement about whether AI systems read structured data in general, and the difference matters - see what this does not show.

Why This Test Had to Be Built

When a pricing page repeats the same value in visible copy and in structured data, a correct answer does not reveal which representation contributed. The observation is unattributable, and any study of such pages inherits that ambiguity whether or not it says so.

To create an attributable comparison without exposing a client site, we built fictional single-layer and conflicting-layer conditions on a domain we control. Those conditions depart from documented Google guidance by design, which is why they are research instruments rather than recommended implementations, and why we say so plainly further down.

One prior test approached it the same way. In October 2025, searchVIU built a single page carrying eight product variants with prices placed in different layers, and found that a price existing only in JSON-LD was returned by none of the five systems tested. Their stated limitation is that the test covers live direct fetch only, on one page, at one point in time. This test repeats that design independently with different pages, a different researcher and a different date.

The Four Pages

Four fictional single-product vendors, one page each, published August 29, 2026 on a domain CoreAEX controls. Every page carries a visible notice that the product does not exist and the page is part of a study. All four are live and can be opened:

PagePrice in visible textPrice in JSON-LDRole
Kirnwold$47$47Control - both layers agree
Tarnwick$63no structured data at allText only
Velquistnone$89Markup only
Modrenna$34$118Divergent - the layers disagree

Every price is distinct and none is close to another, so any returned number is attributable to one page and one layer without interpretation. The four use unrelated invented vendor names rather than variants of one brand, so that an engine cannot blend them.

Each page remained unchanged throughout the observation window. We recorded a SHA-256 hash of each page as served, before the round and again after it. Each page's post-round hash matched its own pre-round hash - four page-level comparisons, all matching. That check exists because a page edited mid-round would silently invalidate every run in it, and nothing in the answers themselves would reveal that.

Markup was generated from a single template, so the pages share structure while differing in source-layer configuration and in page identity, and it was validated against the published Schema.org release file - version 30.0, released March 19, 2026 - with no unknown types, no unknown properties, no domain mismatches and no terms from the vocabulary's pending area. Documented

The Prompts, and the Check That Makes a Null Readable

Each run supplied the page URL directly and asked one question. Every run was a fresh session with no history - a repetition asked in the same thread measures the model's memory of its own previous answer, which is a mistake we made in a first pass and corrected before collecting anything.

Two prompts, in separate sessions:

  • The retrieval check: Open [URL] and tell me what the [named feature] feature does. Each page carries one distinctively named feature in its visible text.
  • The pricing question: Open [URL] and tell me what this page says about pricing.

The pricing question is deliberately not "what is the price." That phrasing presumes a price exists and pushes an engine toward producing a number; on the markup-only page the honest answer is that the page states none, and a leading question would have converted that into fabrication. We would then have measured our own prompt. Recommendation

The access check is what keeps a null result from being uninterpretable, and it is the part most easily skipped. Without it, "the engine did not return the price" and "the engine never reached the page" produce identical silence, and the second is a live possibility for a page published the previous day on a small domain.

It is weaker evidence than it first appears, and the design is why. Running the check in its own session was the fix for a contamination problem - a second question in the same thread can be answered from the first answer rather than from the page. That fix cost per-response confirmation: a successful check shows the engine could reach and read the page around the observation window, not that a different session did. Where a pricing response itself reports distinctive page content, it is content-access-confirmed on its own terms; where it does not, it is access-capable but content-access-unverified, and this page says which.

Engines that failed the access check entirely are excluded from the denominators rather than counted as evidence, and the excluded count is reported alongside every figure below.

Repetitions were allocated where they change what can be said: three each for the two decisive pages, one each for the two controls. Conditions were logged-out sessions from Serbia with browsing enabled, on August 30, 2026. Model versions are unknown - none of these engines exposes a version string to logged-out users - so the finding is anchored to the observation date rather than to named model releases.

Result: A Price That Exists Only in Markup

Velquist carries 89.00 in its JSON-LD and states no price anywhere in its rendered text. Its visible copy says the product is billed per user, per month, in USD, and that there is one plan - but never how much.

EngineRepetition 1Repetition 2Repetition 3Retrieval check
ChatGPTno priceno priceno pricepassed ×3
Geminino priceno priceno pricepassed ×3
Google AI Modeno priceno priceno pricepassed ×3

Across nine pricing prompts, no response returned $89.

The nine matching access-check sessions all succeeded: every engine described the Marrowsync ledger feature accurately from that page. Those checks show the engines could reach and read the page during the window. They are not per-response proof, because each ran in its own session. A pricing response is content-access-confirmed only where that response itself reports distinctive page content; the rest are better described as access-capable but content-access-unverified. Several of the nine meet the stricter bar on their own terms - they describe the page's plan structure and billing wording while stating that no figure appears - and the count below does not depend on which reading is applied.

The engines did not simply omit the number. Several reported its absence directly, in the same answer in which they described the page's contents accurately:

  • ChatGPT: "The page does not state a numeric per-user monthly price. It only says Velquist is billed per user, per month, in USD, and that it has a single plan and price."
  • Google AI Mode, repetition 1, translated from Serbian: "the price for Velquist is not stated … the exact numeric amount is nowhere written in the text of the research page."
  • Google AI Mode, repetition 3: "This page does not contain any concrete price or figure for Velquist."

That last phrasing is the most precise thing any engine said in the round. The response had access to the page's visible content. It reported what the text contained. The served page contained the JSON-LD value; the answer did not reproduce it.

Result: When the Two Layers Disagree

Modrenna shows $34 to a reader and carries 118.00 in its JSON-LD. This is the page that identifies a layer in both directions: whichever number comes back names the layer it came from, so a positive answer is attributable rather than a null needing interpretation.

EngineRepetition 1Repetition 2Repetition 3
ChatGPT$34$34$34
Gemini$34$34no number returned
Google AI Mode$34$34$34

Across nine pricing prompts: eight responses returned the visible $34, one returned no number, and none returned the $118 JSON-LD value. No engine mentioned a second price, and none flagged the discrepancy between what the page displays and what it declares.

The single run that produced no number returned a description of the page instead. It returned no wrong number, and no run anywhere in the cell returned the markup value.

Result: The Controls

The controls exist to show the pipeline works, and one of them turned out to carry a finding of its own.

Kirnwold states $47 in both layers. 2 of 3 engines returned $47; the third returned no content for this page. Gemini reported it could not access this one page while returning accurate content from the other three in the same window, which is an inconsistency we can describe but not explain, and it is recorded rather than smoothed over.

Tarnwick states $63 in visible text and carries no structured data whatsoever - no JSON-LD, no Microdata, no RDFa. 3 of 3 engines returned $63.

That second control is quieter than the headline results and it does real work. Three engines extracted a price accurately from a page with no markup on it at all. Whatever produced the correct answers on the control and the divergent page did not require structured data to be present in order to function.

Two Engines Produced No Data

Claude and Perplexity did not successfully return content from any of the four pages. Claude reported, on every attempt, that the site "blocks automated access via robots.txt, so the fetch was refused before I saw any content." Perplexity reported fetch failures and then sign-in walls.

Our robots.txt does not support that. Re-fetched at the close of the round, in full:

User-agent: *
Disallow: /admin/
Disallow: /includes/
Disallow: /config/
Disallow: /logs/
Disallow: /backup/
Disallow: /.env

Sitemap: https://quietfailurerecords.com/sitemap.xml

Nothing disallows the research directory, and three other interfaces returned accurate content from those URLs in the same window. Anthropic documents Claude-User as the agent used when individuals ask questions to Claude, and states that its crawlers respect 'do not crawl' signals by honoring industry standard directives in robots.txt. Documented

The server access log shows no entries for that fetcher. That narrows the cause without settling it, and the distinction is worth keeping: a request refused by the fetcher on its own robots.txt evaluation and a request dropped by a firewall or CDN before it reached the origin both leave an empty origin log. We report the failure and not a diagnosis.

Neither engine's silence is evidence about markup, and it would be easy and wrong to fold two blank columns into a "none of five systems" figure. Both are excluded from every denominator on this page. This is also why the study runs five engines and reports three: losing two of them to access problems is an ordinary outcome of testing live systems, and pretending otherwise would inflate the sample.

What the Validators Did and Did Not Catch

Before the answer runs, all three markup-bearing pages went through the Schema Markup Validator and Google's Rich Results Test. Two results are worth publishing.

Neither tool flagged either manipulation. Velquist's price being absent from the visible page, and Modrenna's markup contradicting its own displayed price, passed both tools without comment. Google's General Structured Data Guidelines state the rule directly - Don't mark up content that is not visible to readers of the page and Your structured data must be a true representation of the page content - and the tooling does not enforce either sentence. These validators check syntax and field presence, not truthfulness. Documented

The Rich Results Test reported the pages as valid while listing a documented requirement as optional. Google's Software app documentation lists, under required properties, name, offers.price, and one of aggregateRating or review. Our pages carry no rating and no review, by design - inventing them would have put fabricated data on a research page. The Rich Results Test returned "1 valid item detected" and raised the missing rating as a non-critical issue marked optional.

We are recording that as an observation, not a conclusion. One tool, one date, one feature. It does not tell us whether the discrepancy is a severity-labelling choice, a difference between item validity and feature eligibility, or something else. What it does show is a mismatch between the tool's summary and severity presentation and the feature documentation: a "valid item" headline with the missing property marked non-critical can be read as full feature eligibility, while the Software app documentation lists that property as required. Neither validator tests whether a price is truthful or visible with the completeness of a human parity audit.

What This Does Not Show

The temptation with a 0-of-9 result is to widen it into a claim about how AI systems treat structured data. The design does not carry that, and stating the limits precisely is the difference between a useful finding and a quotable one.

Three engines, one date, four pages, one domain. ChatGPT, Gemini and Google AI Mode, on August 30, 2026, from Serbia, logged out. Claude and Perplexity contributed nothing. A different date could produce a different result and we have already seen how quickly these systems move: the same pages were reachable by name in one engine on August 29 and not reachable by name in the same engine the following day.

URL-supplied content access only. Every prompt supplied the URL, but this design does not establish whether an interface used a live request, a search index, a cached representation or another retrieval layer. It also says nothing about training data or third-party copies. searchVIU flagged a comparable boundary on their test.

Model versions are unknown for all three engines, because none exposes one to logged-out users. The result is anchored to a date, not to a named release, and a reader cannot assume the same system generation would behave identically later.

These are fictional products on a small domain with no inbound links and no authority in the software category. That is close to the hardest retrieval case there is, which makes a positive result strong and a null weaker. The nulls here are load-bearing, so this limitation cuts against our own headline and belongs beside it rather than in a footnote.

This is a practitioner test, not peer-reviewed research. It is the second such test we are aware of on this question. Two practitioner tests agreeing is worth more than one, and it is still two practitioner tests.

What the result does support is narrow and actionable: a commercial fact that exists only in markup, and nowhere a reader can see it, has no demonstrated route into an answer in these conditions. That is enough to justify keeping prices, billing periods, seat minimums and commitments in visible copy - a recommendation that would stand on accessibility and reader-trust grounds even if this test had come out the other way.

What We Deliberately Broke

Two of the four pages depart from Google's published guidance, and that departure is the experiment rather than an oversight.

Google's General Structured Data Guidelines state: Don't mark up content that is not visible to readers of the page. Velquist breaks that rule. Modrenna breaks it and also departs from Your structured data must be a true representation of the page content. The same documentation notes that a structured data issue can result in a manual action. Documented

Three things follow, and we did all three:

  • The pages are hosted on a domain that is not CoreAEX and not a client's. Whatever risk attaches to them attaches there.
  • The Manual Actions report was checked before the round and at its close. None at either point. None appeared at either check. Such an action would be material to the validator portion of this study, and Google states its consequence precisely: A structured data manual action means that a page loses eligibility for appearance as a rich result; it doesn't affect how the page ranks in Google web search. It is not described as removing the page from the index. Any action would be recorded as a study event rather than treated as automatic invalidation of the AI-answer observations.
  • Every page carries a visible notice stating that the product is fictional, is not for sale, and that the page belongs to a study.

Do not do this on a production site. The technique here is a measurement instrument, not a tactic, and there is nothing to gain from it commercially - the result is that the hidden value did not travel. Recommendation

Reproducing This

The four pages are live at the URLs above and unchanged. Anyone can open one, ask an engine to read it, and compare. If you do, three things determine whether your result means anything:

  • Give the engine the URL. Asking by product name tests whether a fictional product with no inbound links is discoverable, which is a different question and one these pages mostly fail. We started with name-based prompts and had to discard them for exactly that reason.
  • Use a fresh session for every question. A second question in the same thread can be answered from the first answer rather than from the page.
  • Ask the retrieval-check question first, in its own session. Without it you cannot tell a page that was read and reported accurately from a page that was never opened.

The full run record - every prompt, every verbatim response, the coding rules, the page hashes and the exclusions - is held with this study and available on request. Recommendation

The round repeats on a fixed schedule against the same unchanged pages, with the hashes re-verified each time. Results that move will be published as movements rather than quietly replacing what is here, because a live-systems finding that changes is itself the finding.

Wondering whether your own pricing survives this?

The check is quicker than the study: open your pricing page, and ask whether every commercial fact a buyer needs is readable without a parser. If some of it only exists in markup, that is a short conversation. Book a Session.

Sources

Sources and data: The four test pages are published by Quiet Failure Records. Each was verified unchanged by matching its post-round SHA-256 hash to its own pre-round hash, taken from the served response before and after the observation window: Kirnwold, Tarnwick, Velquist, Modrenna. Both properties are operated by the author, which is disclosed here rather than left to be discovered. Markup was validated against the Schema.org vocabulary release version 30.0 (released March 19, 2026). Google's structured data rules on visibility, truthfulness and manual actions are from General structured data guidelines (last updated July 10, 2026); the Software app required properties are from Software app structured data (last updated December 10, 2025). Anthropic's crawler agents and robots.txt position are from Anthropic's crawler documentation. The prior test using a comparable URL-supplied design is searchVIU, "Schema Markup and AI in 2025" (Michael, published December 2, 2025, tests run October 30, 2025 - a practitioner test on a single page with eight product variants across five systems). Observation runs for this study were carried out on August 30, 2026 from Serbia, logged out, with browsing enabled; model versions are not recorded because none of the tested engines exposes one to logged-out users.


About the author

Zarko Zivkovic is the founder of CoreAEX, building technical SEO, AEO, and AI-visibility systems for B2B SaaS companies. Connect on LinkedIn.