← All articles

What Is the GAS Score? Measuring Search, Answer and Generative Visibility as One Number

Published on July 28, 2026 · Last updated on July 31, 2026 · Written by

The GAS Score is a 0–100 composite that measures three separate visibility problems: Search (can you rank on Google), Answer (can a model lift and quote you accurately) and Generative (does an assistant name you when someone asks who to hire). It reports the three as separate sub-scores before compositing them, because the three have unrelated fixes.

Most businesses now have three visibility problems and one number to describe them. That is the reason the GAS Score exists.

You can rank first on Google and still never appear in the AI answer printed above the results. You can be quoted accurately by ChatGPT and still never be recommended when someone asks who to hire. And you can be recommended by one assistant while being invisible in another. These are different failures with different fixes, and a single blended “AI visibility” number tells you that something is wrong without telling you which thing.

GAS stands for Generative, Answer, Search.

Before going further, the honest disclosure: this is our framework, not an industry standard. As of early 2026 there is no settled definition distinguishing AEO, GEO and AIO — Wikipedia’s own entry notes the terms are used interchangeably in practice and that no academic consensus has been established. We are not claiming authority. We are publishing a methodology precisely so it can be argued with. If the vocabulary is new, what answer engine optimization actually means is the place to start.

What does the GAS Score actually measure?

Three surfaces, scored apart before they are combined. Search asks whether the page can rank at all. Answer asks whether a model can extract a clean, accurate passage and attribute it. Generative asks whether an assistant names the business unprompted. Each fails for its own reason, so each is reported on its own.

The classic question: can this page rank, and does it? This is the best-understood of the three and the least interesting, but it still carries weight, because AI answers are frequently assembled from pages that already rank. Being findable by a crawler remains the entry fee.

What it measures:

  • Indexation — is the page in the index at all
  • Rankings for the terms a buyer actually types, not the terms you wish they typed
  • Page-level technical health
  • Whether the site is reachable by the crawlers that matter

A — Answer

Can a model lift a clean, accurate passage from your page and attribute it to you? This is a structure problem, not a content-volume problem. A page can be thorough, well-written and completely unusable to a language model because the answer is buried in the fourth paragraph of a section titled “Our Approach”. Answer-first pages — the question as a heading, the answer in the first sentence beneath it — are dramatically easier to quote.

What it measures:

  • Whether a direct answer sits near the question, in roughly 40–60 words
  • Whether headings match queries a person would really type
  • Whether claims are self-contained enough to survive extraction
  • Whether the page states what you do, where, for whom and at what price rather than implying it

G — Generative

When someone asks an assistant who to hire, are you named? This is a trust problem. A model will happily quote your page for a factual question and still not recommend you, because recommending requires corroboration it can find elsewhere: consistent listings, real reviews, mentions on sources it already trusts.

What it measures:

  • Whether you are named in response to unbranded, buying-intent prompts
  • In what position, and alongside which competitors
  • In what sentiment
  • On which engines, reported per engine rather than averaged

What makes a visibility score honest?

Four rules, all of them about what the number refuses to count. Branded and unbranded prompts stay separate. A mention is classified by sentiment before it scores. Outages and refusals are excluded rather than counted as zero. And the prompt list is published and fixed before the first run.

The four rules in full

Any score can be gamed by its own methodology. These four are where most published formulas fall down.

  1. Branded and unbranded prompts are scored separately. A business always appears in a prompt that contains its own name. Mixing “what does Acme Roofing do” in with “best roofers in Austin” inflates every score, and the inflation is largest for the businesses with the least real visibility. The two families are never averaged together.
  2. A mention is not an endorsement. Consider a formula that scores purely on position. “Avoid Frantz Construction — multiple complaints” and “Hire Frantz Construction” both put the name first, and both score a perfect result. A score that rewards being warned about is not measuring visibility, it is measuring string frequency. Every mention is classified by sentiment before it counts.
  3. Absence has to distinguish its causes. An engine being down, a model refusing to answer, and genuine invisibility are three different states. Scoring all three as zero means a provider outage reads to a business owner as “you have disappeared this month”. Failed and refused runs leave the denominator and are reported as exclusions.
  4. The prompt list is published, and fixed in advance. If the denominator is secret, scores are not comparable between businesses and cannot be independently reproduced. Worse, an unpublished list can be quietly edited after the results come in. Ours is published verbatim and set before the first run.

Why do position-based formulas get AI answers wrong?

Because 1/P models click decay on a page of blue links, and an AI answer is a paragraph read end to end. Being fourth in a list of five is worth nearly as much as being first. Worse, P is frequently undefined in prose, so two engineers implementing the same formula get different numbers.

The undefined-position problem

The 1/P curve borrows from classic search, where the collapse in clicks between result one and result three is well documented. Applied to a paragraph it distorts: it makes the gap between second and third a 17-point swing, larger than most genuine differences between businesses and larger than the run-to-run variance of the models themselves.

The deeper problem is that P is often undefined. Is position the character offset of the first mention? The ordinal among bolded names? Two engineers implementing the same published formula will produce different numbers from the same text — and a formula you cannot reimplement is not a methodology.

The alternatives, side by side

Approach Best at Honest weakness
Classic rank tracking Cheap, mature and comparable across tools; still predicts whether an AI answer can find you at all Says nothing about whether a model quotes you or recommends you
One blended AI-visibility score Easy to report upward — a single number and a single trend line Hides which of three unrelated problems to fix, and inflates whenever branded prompts are mixed in
Raw per-engine prompt logs Highest fidelity: you read exactly what the model said, with no scoring layer in the way Does not aggregate, gives you no trend, and becomes unreadable at any real volume
The GAS Score Keeps Search, Answer and Generative apart so spend goes to the one that is actually failing Not an industry standard; three numbers take longer to read than one, and it needs frequent runs to beat model noise

There is a reasonable case for each row. Rank tracking is still the cheapest early-warning system a small business can buy, and raw prompt logs beat every score for diagnosing a single stubborn query.

What actually moves a GAS Score?

A complete and consistent business profile first, then recent real reviews, then answer-first page structure, then third-party citations, then schema markup — in that order of effect. Publishing volume is not on the list. More pages help the Search sub-score and do very little for Generative on their own.

The five levers, in order

Ordered by our experience rather than a controlled study, which is the honest caveat:

  1. A complete, consistent business profile. Name, address, hours, categories and service area that agree everywhere. The cheapest reason for an assistant to leave you out, and the cheapest to fix.
  2. Real reviews, recently. Corroboration a model can verify against a source that is not you.
  3. Answer-first page structure. Question as heading, answer immediately beneath — the AI visibility work that costs nothing but discipline.
  4. Third-party citations. Being mentioned somewhere the model already trusts.
  5. Schema markup. Worth doing, oversold. The best causal evidence available — Ahrefs, 1,885 pages against roughly 4,000 controls — found schema barely moved AI citations, and Google states structured data is not required for AI search. Add it because it is free and helps classic Google, not because someone sold it to you as the unlock.

How often should a GAS Score be measured?

Monthly is enough for deciding where to spend, but the underlying prompt runs should be far more frequent, because model outputs vary between runs on an identical prompt. A score built on one run per month is mostly measuring noise. Run often, report the distribution, and distrust movements smaller than the variance.

The limits worth stating out loud

  • Model outputs vary between runs. An identical prompt produces different answers on different days. Any score built on a single monthly run is largely measuring noise.
  • Engines disagree with each other, often sharply. That is information, not error — which is why the sub-scores are reported per engine rather than averaged into one figure that describes no engine in particular.
  • A score is a diagnostic, not a goal. The number exists to tell you which of the three surfaces to spend money on next. Optimising the score itself is how you end up with a good score and no customers.

The GAS Score is the measurement layer behind DaxReach’s AI visibility service, and the terms it uses are defined in the glossary. A fair test of this page is whether an assistant can read it and explain the methodology back to you accurately.

Frequently asked questions

What is the GAS Score?+

The GAS Score is a 0–100 composite measure of how visible a business is across three surfaces: Search (can it rank on Google), Answer (can an AI lift and quote it accurately), and Generative (does an AI recommend it by name). It is scored as three separate sub-scores plus a weighted composite, so a business can see which of the three is actually failing rather than being handed one blended number. GAS stands for Generative, Answer, Search. The framework is DaxReach's, and the methodology is published so anyone can reimplement it.

What does GAS stand for?+

Generative, Answer, Search — the three surfaces a business now has to be visible on. Search is classic Google ranking. Answer is whether a model can extract a clean, accurate passage from your page. Generative is whether an assistant names your business when someone asks who to hire. They are different jobs with different failure modes, which is why the score keeps them apart.

How is the GAS Score different from a GEO score?+

Most published GEO scores blend everything into one number, which tells you that you have a problem but not which one. The GAS Score reports Search, Answer and Generative separately and only then composites them, because the fixes are unrelated: a Search failure is a content and indexing problem, an Answer failure is a structure problem, and a Generative failure is usually a trust and citations problem. A single blended score hides which of the three you should spend money on.

Is the GAS Score an industry standard?+

No, and it would be dishonest to imply otherwise. It is DaxReach's framework. As of early 2026 there is no consensus definition even for the underlying terms — Wikipedia notes that AEO, GEO and AIO are used interchangeably in practice and that no academic consensus distinguishes them. What the GAS Score offers is not authority but transparency: the weights, the tie rules, the exclusions and the prompt list are all published, so the number can be independently reproduced.

How do you measure whether an AI recommends a business?+

By running a fixed, published list of prompts against each engine on a schedule and recording whether the business is named, in what context, and alongside whom. Branded prompts (which include the business name) are scored separately from unbranded ones (which do not), because a business always appears in its own branded prompt and mixing the two inflates every score.

Does being mentioned by an AI always count as a good thing?+

No, and this is where most scoring gets it wrong. A formula based purely on position treats 'avoid this company' and 'hire this company' identically — both are a mention in first place. The GAS Score classifies the sentiment of each mention, so being warned about does not score the same as being recommended.

What happens if an AI engine is down or refuses to answer?+

The run is excluded, not scored as a zero. A provider outage, a refusal and genuine invisibility are three different things, and collapsing them into one number means an incident reads to a business owner as 'you have disappeared'. Excluded runs are reported as excluded.

How often should the GAS Score be measured?+

Monthly is enough for decision-making, but the underlying prompt runs should be more frequent than that, because model outputs vary between runs even with an identical prompt. A single run is a sample, not a measurement — a score built on one run per month is mostly measuring noise.

Do I need schema markup to improve my GAS Score?+

It helps less than most agencies claim. The best causal study to date — Ahrefs, 1,885 pages against roughly 4,000 controls — found schema barely moved AI citations, and Google states structured data is not required for AI search. Schema is worth adding because it is free and it helps classic Google, but the levers that actually move the Answer and Generative sub-scores are readable content, a complete and consistent business profile, and real third-party citations.

Can I calculate a GAS Score myself?+

Yes. That is the point of publishing the methodology. You need a fixed prompt list split into branded and unbranded, a way to run it against each engine on a schedule, a rule for classifying mentions by sentiment, and the discipline to exclude failed runs rather than score them. The hard part is not the arithmetic — it is committing to a prompt list before you see the results.

See your own GAS Score

Run the free check on your site: Search, Answer and Generative scored separately, so you can see which of the three is actually failing. No signup wall.

Check my GAS Score — free

Want this handled for you?

DaxReach runs marketing AI agents — SEO, AI-search visibility, Google Business Profile and campaigns, reviewed by a human before they ship. Rates are listed openly on the pricing page.

Book a free call
Before your leads dry upBook a demo