Guide · 2026-08-20

Check your AI visibility.

You do not need a vendor, a subscription or a dashboard to find out how ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews describe your business. You need ten questions, five browser tabs and somewhere to write down what comes back.

This is close to the method we run on the brands we track publicly, minus the automation. Our own frozen set is ten questions too, but where this guide spends its last two on questions about your own business, ours spends those two slots on a head-to-head comparison and a competitor-alternative question instead. It takes an afternoon the first time and about an hour every month after that. Everything it produces is yours, including the parts that are uncomfortable to read.

What you’re actually measuring.

When a customer asks an assistant about your category, the answer either contains you or it does not. That is the whole measurement. It is worth splitting into four questions, because they fail for different reasons and they get fixed in different ways.

Whether you appear at all.

Ask something a customer would ask, without using your name, and see whether your business is in the answer. The name-free part is the whole test. If you ask “what does [your business] do,” the engine will describe you whether or not it would ever have brought you up on its own. That answer is worth having, but it measures accuracy, not visibility. Keep the two apart.

What gets said about you.

Appearances are not equal. A recommendation, a neutral namecheck, a warning and a confident description with the facts wrong are four different results. Being named as the expensive choice for a job you stopped doing in 2023 is worse than not appearing at all, because it repeats, and every reader of it decides against you before you know they exist.

How prominently.

Named first in a list of three is not the same result as thirteenth in a list of fifteen, so record where you landed and how long the list was. One exception. If you asked “[you] vs [competitor],” your position in that answer means nothing, because you put yourself there. In those answers, read the description and ignore the order.

How much of the room you’re sharing.

Read the “X of Y” you already wrote down for position, but this time read the Y. An answer that names only you and one competitor is a different contest than one that names six businesses in the same breath. How much of that space is you and how much is competitors is a fourth number worth tracking, alongside whether you appear, what gets said, and where you land.

Run your own check.

Write ten questions.

Ask what a customer would type, not what you would type. The test for a good question is that a stranger with your problem could have written it without knowing your business exists. Eight of the ten must not contain your name. The last two can, and they measure something different.

Ten is the number we use per engine in our own scans. One answer is an anecdote. Ten is a shape you can read.

  1. 1best [category] in [your city]
  2. 2best [category] for [the customer you want]
  3. 3who should I use for [the problem you solve]
  4. 4alternatives to [a competitor]
  5. 5[competitor A] vs [competitor B]
  6. 6how much does [what you sell] cost in [your city]
  7. 7what should I look for when choosing a [category]
  8. 8is [what you sell] worth it for [customer type]
  9. 9is [your business] any good
  10. 10what does [your business] do

Ask them the way people actually buy from you. If your customers care about price, put price in. If they buy locally, put the city in. Generic category questions produce generic answers, and in a generic answer everyone in your category looks equally invisible.

Ask five engines.

ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews. The free tier of each is enough. Google AI Overviews is the summary at the top of an ordinary Google results page, and not every question produces one, so note when it does not rather than leaving the row blank.

Ten questions across five engines is 50 answers. Budget two hours for the first pass. That is a longer commitment than the twenty-minute version on our homepage checklist, which only asks three questions. This is the fuller pass, and it is the one worth repeating monthly as a baseline.

It is worth knowing what you are competing for. When Google shows an AI summary, only 1% of searches end in a click on a source cited inside it, from roughly 69,000 searches by 900 U.S. adults. Being described inside the answer is now the part that reaches the customer. The click is the bonus.

Pew Research Center, Jul 2025

Keep the conditions clean.

Four rules, and they matter more than the questions do.

  • Start a new chat for every question. Assistants remember the conversation, and an engine that read your website two messages ago is not answering the way a stranger’s would.
  • Log out, or use the temporary chat mode where the engine offers one. Personalization and saved memory can bend the result in your favor, which is the one direction you cannot afford.
  • Do not argue with the answer, and do not follow up with “what about [your business].” Record the first answer, in full, including the parts you disagree with.
  • Run at least one question twice. The same question does not always produce the same answer. That variation is part of what you are measuring, and it is the reason ten questions across five engines beats one question you liked the answer to.

Record it as you go.

Not afterwards, and not from memory. A spreadsheet or a notebook is fine. Part five has a layout that works, and the only column people regret leaving out is the date.

Read your results.

A strong result looks like this. You are named, without being asked for, in at least three of the five engines. You are described in words you would have chosen yourself. You are in the first few options rather than the last. The facts are right. And the engines broadly agree, which is not one of the four things the published score measures, but is still worth noticing.

Weak results come in three kinds, and they are not variations of the same problem.

Absent.

You appear in none of the eight questions that did not name you. This is the least alarming of the three. Nothing in the answer is wrong about you. There is simply nothing to retrieve, or nothing that says clearly enough that you do this specific thing, for these people, in this place.

Described wrongly.

You appear, and the facts are off. Old prices, a service you dropped, the wrong service area, or a confident paragraph about a different business with a similar name. This is the worst of the three, because it converts nobody and it travels. Nothing corrects it on its own, and the version of it living on somebody else’s page will outlive the version on yours.

Described well but rarely.

One engine of five, or two questions of ten, and the answer is accurate when it comes. Something exists to retrieve and nothing corroborates it. This one flatters and misleads in equal measure, because the answer you found reads well enough to feel like a result.

Mentions and citations are different things.

A mention is your name in the text of the answer. A citation is a link the engine hangs a fact on. You can have either without the other, and they get fixed differently, so give them separate columns. The list of domains an engine cites around your category is a map of where the answer comes from, and it is often the most useful thing on the page.

What this method cannot tell you.

  • What any of it is worth. Fifty answers describe how you are being represented. They say nothing about traffic or revenue, and anyone converting one into the other is guessing.
  • Whether you caused a change. Engines get retrained, the sources they pull from when answering change, and the same question varies between runs. A movement between two months is a signal worth watching, not proof of anything.
  • What the model learned during training versus what it looked up while answering. From the outside those are indistinguishable, and they move on completely different timescales. Your site can change the second within weeks. The first only changes when a model is next trained.
  • Whether being read means being recommended. As of June 2026, 52% of AI crawler requests were made for model training rather than to answer a live question. Being crawled constantly and named never is an ordinary result.
  • Anything with a margin of error. Fifty answers collected in one afternoon by someone with a stake in the outcome is a small sample read by an interested party. Treat a single question changing as noise. Look at the shape across all ten.

Cloudflare, Jul 2026

What to fix first.

Each result below points at a different item on the eight-item checklist on the homepage. The numbers refer to that list. Find the line that describes what you found and start there, rather than starting at the top of the checklist and working down.

Absent from every engine.

Check whether you are blocked before you write a word. Open yoursite.com/robots.txt and look for GPTBot, ClaudeBot, PerplexityBot and Google-Extended. If those are disallowed, nothing else on the list can work. That is item 1. Once you are readable, the work is items 3 and 4: make each heading the question a customer actually asks, answer it in the two sentences underneath, and write your prices, coverage and hours in plain text rather than in an image, a PDF or a widget.

Present in one or two engines, absent in the rest.

One likely explanation: the engines that miss you are reading sources that do not mention you, in which case the fix is not on your website. Item 7 is the lever: listings, directories, forums and reference pages. Correcting how you are described on one widely-read third-party page can move more than a month of work on your own. Item 6, structured data, makes the facts you do control harder to misread. But check item 1 first. One engine going quiet while the others find you is also the signature of a crawler-specific robots.txt rule, and that is a five-minute fix, not a content project.

Present, but the facts are wrong.

Start with your own pages, item 4, because that is the copy every engine can reach fastest. Then find the third-party pages repeating the same error and fix those, item 7. Then date everything, item 5. An engine choosing between two claims that contradict each other has very little to go on beyond which one looks current.

Present, but described vaguely.

“A local option worth considering” means the engines found your name and nothing else. That is item 4. Say who it is for, who it is not for, what it costs and where you work. Vague pages produce vague sentences, and a vague sentence never wins a recommendation.

Present, but never near the top.

This tends to be the slowest of these to move, because it depends on more than your own pages. Items 3, 6 and 7 together, then measure again next month. There is no fast fix here, and anyone offering you one is offering it to someone who has no baseline to check it against.

Nothing to compare against.

If this is your first run, that is the finding, and it is worth having. Item 8 is the whole point of part five below. A single run tells you where you stand. Two runs tell you which direction you are moving.

Keep a baseline.

A month from now, the value of today’s answers is the comparison, not the answers. This is the step that turns an interesting afternoon into a measurement, and it is the one people skip.

One row per engine per question. Fifty rows a month, in whatever you already use.

date        engine               question                   appeared  position  described as          cited
2026-08-20  ChatGPT              best [category] in [city]  yes       2 of 6    "good budget option"  yelp, reddit
2026-08-20  Claude               best [category] in [city]  no        n/a       n/a                   none
2026-08-20  Gemini               best [category] in [city]  yes       5 of 9    "family-run"          your site
2026-08-20  Perplexity           best [category] in [city]  no        n/a       n/a                   none
2026-08-20  Google AI Overviews  best [category] in [city]  no        n/a       n/a                   none
  • Date every run and keep the old ones. Never overwrite a row. The old rows are the entire asset.
  • Do not change the questions. We freeze the question set for the brands we track publicly, for exactly this reason. Change the question and the number stops meaning anything. If you need a new one, add it as a new question and leave the original running.
  • Keep the cadence steady. Monthly is enough for almost every business, and it is slow enough that you will actually do it.
  • Record which engine, never just “AI.” They diverge from each other, and the divergence is where most of the useful signal lives.
  • Write down what you changed on your side, dated, in the same file. Six months of results with no record of what you did is a chart nobody can read.

If you want to audit our numbers.

We publish an AEO Score for the brands we track. It is the four things this guide asks you to write down, weighted and rolled into one number between 0 and 100. A score is only worth anything if a stranger can check it, so the formula is below rather than described.

You do not need this to run the method above. It is here for anyone who wants to reproduce a figure we have published.

How we score it

Formula

AEO Score = presence        × 0.35
          + quality         × 0.25
          + prominence      × 0.20
          + share of voice  × 0.20

Each component is scored 0 to 100 before weighting. The composite is rounded to a whole number.

Components

Presence 35%

Whether you appear at all, in questions that never named you.

Quality 25%

Whether the mention is a recommendation, a namecheck, or a warning. Scored on every question, including the two branded ones.

Prominence 20%

Where you land in the answer, counted only for the questions that never named you.

Share of voice 20%

Of the organic answers that named a brand at all, yours or a competitor’s, how often it was you.

What the score excludes

  • Questions that named the brand count nothing toward Presence. Asking an engine about a brand and then scoring it for answering is circular.
  • Comparison queries do not count toward Prominence at all, for the same reason part one gives: landing first in an answer you asked for yourself is not a result.
  • Quality is the exception to that pattern. It is scored on every question, including the two branded ones, because branded questions tend to draw out the most opinionated answers.
  • Presence carries the heaviest weight, because showing up at all is the precondition for everything else in the formula.
  • Prominence and Quality are calculated only from queries where you were mentioned, so we scale both down when Presence is thin. Full credit applies once Presence reaches 25% or above, ramping to zero below that, so one lucky mention cannot outscore a brand that shows up everywhere.
  • We used to score a fifth dimension here, agreement across the five engines. We do not anymore. Engines are still scanned to gather evidence, but the score reflects performance across the whole query set, not how many engines happen to agree in one scan.

The component definitions, the score bands and how we classify a mention are on the methodology page.