Check whether AI search can read your site
More and more buyers ask an assistant before they ask a shop. That makes an uncomfortable question worth answering with evidence: when a machine reads your website, what does it actually find?
Most answers to that question are one-off opinions. Someone runs your homepage through a chatbot, gets back a confident-sounding review, and you have no way to tell which parts are true, no way to repeat it next month, and no way to compare yourself with anyone else.
The AI Search Review does the opposite. It fetches your pages the way a crawler would, checks a fixed list of things, and gives you the same numbers every time. Run it on your own site, on a competitor, or on a peer you admire — the answer is a measurement, not a guess.
What it looks at
Section titled “What it looks at”A page has to clear four stages before an assistant can answer a question from it. The review reports on each:
| Stage | The question | What it checks |
|---|---|---|
| Reachable | Can a crawler get in at all? | robots.txt rules, crawl delays, whether your sitemap is announced |
| Retrievable | Can it find your pages? | Sitemaps, how many pages are listed, redirects and canonicals |
| Extractable | Can it read your product facts? | Product structured data and which properties it carries |
| Answerable | Is the answer actually there? | Specification tables, page depth, description quality |
It uses no AI and no browser. Every finding is a fact about what your site served on the day of the run, which is exactly what lets you re-run it later and compare.
Run it on one site
Section titled “Run it on one site”Add an AI Search Review step to a flow and give it a URL.
| Field | What to put |
|---|---|
| Site URL | https://www.example.com — the site to review |
| Products to sample | How many product pages to fetch. Defaults to 25; 0 means every page the sitemap lists |
| Timeout | Seconds to wait per request. Defaults to 15 |
Only public pages are fetched, only with ordinary GET requests, and only pages the site lists in its own sitemaps. Nothing is submitted, logged into, or changed.
The step gives you back:
report_md— a readable report you can drop straight into a documentfindings— each issue with a severity and the evidence behind itsummary— headline counts (pages listed, pages checked, findings by severity)coverage— the share of sampled pages carrying each product propertycrawlers— which crawlers are allowed or blocked, and what each one is forpages— the per-page detail, if you want to dig
A worked example
Section titled “A worked example”Point it at your own shop with a small sample first:
- Site URL →
https://www.example.com - Products to sample →
10
A typical result:
[HIGH] robots.txt asks every crawler to wait 60s between pages At this rate a full pass over the catalogue takes hours.
[MEDIUM] robots.txt does not point to the sitemap A crawler that starts at robots.txt never finds the sitemap index.
[MEDIUM] No specification table on any of 10 sampled pages Facts live in prose only.
Each finding names what was measured and how many pages it applied to, so you can check it yourself.
Why “which crawler” matters more than “how many”
Section titled “Why “which crawler” matters more than “how many””The review reports crawler access split by purpose, and this is the part most advice gets wrong.
- Search and answer crawlers —
OAI-SearchBot,PerplexityBot,Claude-SearchBot,Googlebot,Bingbot. These fetch pages in order to answer people’s questions. Blocking one removes you from that surface. OpenAI says so plainly: sites opted out ofOAI-SearchBotwill not appear in ChatGPT’s search answers. - Training crawlers —
GPTBot,ClaudeBot,CCBotand others. These gather material for training models. Blocking them costs you nothing in assistant answers, and plenty of publishers do. - User-triggered fetchers —
ChatGPT-Userand friends, which fetch a page because a person asked about that specific link.
Treating all of these as one category is how sites accidentally remove themselves from AI answers while trying to opt out of training — or leave a crawl delay in place that quietly limits how much of the catalogue ever gets read.
Compare yourself with the field
Section titled “Compare yourself with the field”Because the URL is asked for at run time rather than configured, the same review runs against anyone. Run it on your own site and on two or three competitors in one sitting, and the differences that matter show up immediately — as will the gaps that are common to everyone in your category, which are usually the more interesting opportunity.
Keep competitor samples small and don’t put them on a schedule. Reading a few public pages is ordinary behaviour; hammering someone else’s server is not.
Track it over time
Section titled “Track it over time”Write each run into a table with a Create record step and you turn a one-off audit into a baseline. The second run becomes a comparison rather than another opinion — which is what tells you whether a fix actually landed.
Pair it with a Schedule trigger to re-check your own site monthly. A theme update or a plugin change can quietly alter what crawlers see, and this is how you find out in a week rather than a quarter.
What it can and cannot tell you
Section titled “What it can and cannot tell you”It can tell you what is true of your pages today: what a crawler can reach, what facts are readable, where the answers buyers want are missing, and how that compares to other sites.
It cannot tell you whether any AI system will rank, cite or recommend you. That is not measurable from outside, and the review does not claim it. Treat structured-data findings as housekeeping that helps machine consumers such as product feeds and rich results — not as a lever on whether an assistant mentions your brand.
Where to go next
Section titled “Where to go next”- HTTP Fetch — fetch a single page yourself when you want the raw body
- Set up and activate an installed Job — arm the schedule once you’re happy with the results
- Working with data: tables & imports — where to land each run so you can compare them