Does Schema Markup Help AI Citations? What the Evidence Actually Shows
Barely. The best causal study to date — Ahrefs, 1,885 pages that added schema against roughly 4,000 matched controls — found AI citations did not rise, and slipped 4.6% in AI Overviews. Google states structured data is not required for its AI features. Add schema because it is free and helps classic Search, not because someone sold it as the unlock.
A roofing contractor pays an agency $2,400 for an “AI search readiness” package. The deliverable arrives 6 days later: structured data across 40 pages, a validator screenshot with 0 errors, and a report calling the site AI-ready. Four months on, asking ChatGPT who to hire for roof replacement in that metro still returns 3 competitors and no mention of the client.
Nothing was broken. The markup is valid, the pages render, the validator is genuinely green. The work simply had almost nothing to do with the outcome it was sold against — and the reason is measurable rather than a matter of opinion.
Does schema markup help AI citations?
Not meaningfully. Ahrefs tracked 1,885 pages that added structured data against roughly 4,000 matched control pages and found citations barely moved — down 4.6% in AI Overviews. Google states structured data is not required for its AI features. Schema earns rich results in classic Search; it does not buy a mention.
The study design that makes it credible
Most claims about AI search are correlational: someone notices that cited pages tend to carry structured data and concludes the markup caused the citation. That inference fails because pages carrying good schema also tend to be on larger, better-maintained, more-linked sites — the markup is a symptom of a competent site, not the cause of the citation.
The Ahrefs work matters because it used matched controls. Pages that added schema were compared against roughly 4,000 comparable pages that did not, which isolates the change from everything else that differs between sites. That design is what turns “cited pages have schema” into “adding schema did not produce citations.”
A 4.6% decline is best read as noise around 0 rather than as evidence that schema hurts. The honest conclusion is not that structured data is harmful. It is that the effect size is indistinguishable from nothing.
Why do agencies still sell schema as the AI unlock?
Because it is the most sellable deliverable in the category. Schema is a one-time technical change, shippable in under 1 day, and it produces a green validator screenshot that reads as proof. The 4 levers that actually move citations are slower, need editorial judgement, and resist being packaged into a fixed-scope invoice.
The incentive, traced
Consider what each option looks like on a proposal:
- Schema markup — fixed scope, 1 developer, delivered in days, verifiable with a free validator.
- Answer restructuring — requires rewriting pages someone already approved, and the client must review every change.
- Entity consistency — needs access to profiles the agency does not control, and 3 rounds of the client finding old listings.
- Third-party corroboration — cannot be guaranteed at all, because it depends on other people publishing.
Only the first can be promised with a date. That is the whole mechanism. It is not usually dishonesty so much as an agency shaping its offer around what it can commit to, and a buyer preferring the option with a delivery date attached.
The tell is the pitch itself. Anyone selling structured data as the thing that gets you into ChatGPT has not read the evidence — a position DaxReach states publicly on its comparison page and in the GAS Engine Playbook.
What actually moves AI citations?
Four levers, in strict dependency order: crawler access, because nothing downstream matters if the page cannot be read; extractable answers of roughly 40 to 60 words under a question-shaped heading; entity consistency across every profile an engine can find; and third-party corroboration from sources you do not own.
The 4 levers in detail
1. Crawler access. Check robots.txt for the 9 user agents that matter — OAI-SearchBot, ChatGPT-User, GPTBot, PerplexityBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot and Bytespider. A site that disallowed these in 2023 to stay out of training data also removed itself from recommendation. This is usually a 1-line fix.
2. Extractable answers. An engine lifts a self-contained claim, not a page. Answers running 150 words past the heading get truncated or skipped in favour of a competitor who said the same thing in 50. The mechanics are covered in what answer engine optimization is.
3. Entity consistency. Name, category, address and phone must match across your site and every profile. Where they conflict, an engine either picks one or declines to name you. This is unglamorous and it is where most local businesses lose.
4. Third-party corroboration. What independent sources say about you carries more weight than what you say about yourself, which is precisely why it cannot be bought as a deliverable. It accrues.
Schema sits below all 4. It is worth having and it is not a lever.
Which lever should you buy first?
Whichever one is currently failing, tested top-down and stopped at the first failure. There is no value in rewriting answer blocks on a site whose crawlers are disallowed, and no value in chasing citations while 2 profiles disagree about your business category. Diagnose in dependency order, then spend.
The levers, side by side
| Lever | Effort | Time to effect | Effect on AI citations | Best at | Weakest at |
|---|---|---|---|---|---|
| Schema markup | 1 day | Days | Negligible — 4.6% down in the controlled study | Rich results in classic Search | The thing it is usually sold for |
| Crawler access | 1 hour | Days | Binary — gates everything else | Cheapest possible fix | Does nothing once already correct |
| Extractable answers | Weeks | 2 to 6 weeks | Large on pages you already rank for | Fast wins on existing traffic | Needs content worth quoting |
| Entity consistency | Weeks | 1 to 3 months | Large for local and service businesses | Fixes “which business is this” | Tedious, spread across accounts |
| Third-party corroboration | Months | 2 to 3 quarters | Largest and most durable | Cannot be copied by competitors | Cannot be guaranteed or scheduled |
Read the “weakest at” column honestly: crawler access is nearly free and often already fine, so it is not a project. Extractable answers are the usual best first spend for a site with existing rankings, and useless for a site with 4 thin pages.
The 3 routes to getting it done
| In-house | Schema-first SEO agency | Human-in-the-loop agents (DaxReach) | |
|---|---|---|---|
| Speed to first change | Slow — competes with other work | Fast on the technical change | Fast, once diagnosis is done |
| Cost shape | Staff time, no invoice | Fixed project fee | Retainer — see pricing |
| Who operates it | Your team | Agency, then handover | Agents run it, humans approve every change |
| Who owns the accounts | You | Often the agency | You, always |
| Effect on AI citations | Depends entirely on discipline | Negligible if schema is the scope | Works the 4 levers in order |
| Best at | Judgement about your own market | Classic-Search technical hygiene | Consistency over months |
| Weakest at | Consistency once quarter-end hits | The generative outcome it sells | Needs a real content surface to work with |
The in-house column is not a strawman. A business with an in-house marketer who reads the AEO versus SEO breakdown and works the 4 levers in order will beat any outsourced arrangement, because they know the market better than a vendor can. The constraint is almost never capability. It is whether the work survives contact with a busy quarter.
Should you add schema anyway?
Yes — just stop paying a premium for it. Structured data still drives rich results in classic Google Search, most content systems emit it automatically, and the marginal cost after the template is written is 0. Keep it, confirm it parses, and move the remaining budget to access, answers, entities and corroboration.
The classic-Search case for keeping it
Schema is infrastructure, and infrastructure is judged by whether it works, not by whether it is exciting. Review stars, breadcrumbs and product details in ordinary blue-link results still influence click-through on the traffic that pays the bills today. That case stands entirely on its own merits and needs no generative story attached to it.
What should change is the budget line it sits on. Schema belongs in technical SEO, priced accordingly, next to sitemaps and canonical tags. It does not belong on an AI-visibility invoice, and a vendor who puts it there is either not reading the evidence or is counting on you not to.
The full vocabulary is in the glossary, the scoring method behind these signals is in what the GAS score measures, and how this work is actually run is on how it works and AI Visibility.
Schema markup is a reasonable thing to have and a poor thing to buy. The 4 levers above are the ones that decide whether an engine says your name.
Frequently asked questions
Does schema markup help AI citations?+
Not meaningfully. The best causal study to date, run by Ahrefs, tracked 1,885 pages that added structured data against roughly 4,000 matched control pages and found AI citations barely moved. In AI Overviews they fell 4.6%. Google states plainly that structured data is not required for its AI features. Schema remains worth adding because it is free and it earns rich results in classic Search, but it is not the mechanism that gets a business named inside a generated answer.
Why do so many agencies sell schema markup as the AI unlock?+
Because it is the easiest AI-search deliverable to sell and to invoice. Schema is a one-time technical change, it can be shipped in under 1 day, and it produces a screenshot of a passing validator that looks like proof. The 4 levers that genuinely move AI citations are all slower and harder to package, so the incentive runs toward the deliverable that demonstrates activity rather than the one that changes the outcome.
What actually moves AI citations?+
Four things, in this order. First, crawler access, since an engine cannot cite a page it is disallowed from reading. Second, extractable answers of roughly 40 to 60 words placed directly under the question a buyer would ask. Third, entity consistency, meaning your name, category and location match across every profile an engine can find. Fourth, third-party corroboration from sources you do not own. Schema sits well below all 4.
Should I remove schema markup from my site?+
No. The evidence says schema does not lift AI citations, not that it harms your site. Structured data still earns rich results in classic Google Search, it costs nothing once a template renders it, and most content systems emit it automatically. Keep it, verify it parses, and stop paying a premium for it as an AI deliverable. Removing working markup would trade a real classic-Search benefit for no gain at all.
Does FAQ schema get me into AI answers?+
Not on its own. FAQ structured data describes content an engine can already read in the visible HTML, so it adds machine-readability rather than persuasion. What tends to earn the citation is the shape of the answer itself, roughly 40 to 60 words, self-contained, and placed immediately under a question-phrased heading. If the visible answer is 200 words of preamble, wrapping it in FAQ markup does not make it liftable.
How long does it take to get cited by AI assistants?+
Structural work lands fast and reputation does not. Crawler permissions and answer restructuring can be recrawled and reflected within days on a site that is already indexed, because you are not displacing anyone from a position. Being named as a recommendation takes considerably longer, since engines weigh entity signals that accumulate over months across independent sources. Expect access and extraction fixes in weeks, and recommendation over 2 to 3 quarters.
Which AI crawlers should I allow in robots.txt?+
The ones belonging to systems you want to be cited by. There are 9 that matter: OAI-SearchBot, ChatGPT-User and GPTBot for OpenAI, PerplexityBot for Perplexity, ClaudeBot for Anthropic, Google-Extended for Gemini and AI Overviews, Applebot-Extended for Apple Intelligence, plus CCBot and Bytespider. Many sites disallowed these in 2023 and 2024 to stay out of training data, then discovered the same rules also excluded them from being recommended.
Can I check whether AI assistants mention my business?+
Yes, and it costs nothing to start. Ask ChatGPT, Gemini and Perplexity the same buying question twice. Once naming your business, which tests whether the engine knows you and describes you accurately. Once without your name, which tests whether it recommends you unprompted. Run each engine separately rather than averaging them, because the 3 disagree often enough that a blended score hides the finding you needed.
Is schema markup still worth doing for classic SEO?+
Yes. Structured data is how classic Google Search builds rich results such as review stars, breadcrumbs and product details, and those affect click-through rate on the ordinary blue-link results that still drive most commercial traffic. The honest framing is that schema is classic-Search infrastructure with no meaningful generative upside. Add it once in your page template, confirm it parses, and spend the remaining budget on the 4 levers that move citations.
What is the difference between AEO and schema markup?+
Schema markup is one narrow tactic. AEO, or answer engine optimization, is the whole discipline of making content extractable and quotable by an answer engine. Schema describes your page to a machine in a structured vocabulary. AEO changes what the page actually says and how it is arranged, so a specific claim can be lifted into a generated answer. Schema is a label on the box; AEO is what is inside it.
Want this handled for you?
DaxReach runs marketing AI agents — SEO, AI-search visibility, Google Business Profile and campaigns, reviewed by a human before they ship. Rates are listed openly on the pricing page.
Book a free call