Fill in before sending
[site URL][Search Console property, or "none"][5 to 10 page URLs]
Check whether my site [site URL] can serve content to AI search such as ChatGPT, Claude, Perplexity, Gemini and Copilot. My Search Console property is [Search Console property, or "none"]. My most important pages are: [5 to 10 page URLs]. This is a read-only audit, do not change anything.
1. robots.txt. Read [site URL]/robots.txt with site_read_url and note the status code. If the file returns 5xx or loops through redirects, say at the top that crawlers may not crawl the site at all. Then work out which rule group applies to each of these crawlers and whether my key pages are allowed: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Googlebot, Google-Extended, Bingbot, Applebot, Applebot-Extended, CCBot, meta-externalagent. When a crawler has a group under its own name, only that group applies and the `User-agent: *` rules do not, so account for that.
2. Sort the crawlers by role. The ones that crawl for search and citations are OAI-SearchBot, Claude-SearchBot, PerplexityBot, Bingbot and Googlebot. The ones that fetch a page live when a user asks a question are ChatGPT-User, Claude-User and Perplexity-User. The ones that collect training data are GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot and meta-externalagent. Flag any blocked search crawler at the highest severity, because the site cannot be cited in that assistant's answers. Do not treat a blocked training crawler as an error, describe it as a choice. Note that Google-Extended does not affect Google Search or AI Overviews and only controls whether the content is used for Gemini models.
3. Status codes and headers. Read each key page with site_read_url and max_bytes=0. Report only the status code, the redirect chain (redirects), the `x-robots-tag` and `content-type` headers and the `server` value, not the full header dump. Flag redirect chains longer than two hops, a final URL that differs from the canonical, and any noindex or nosnippet header. If a page returns 403, 429 or 503 and the server value shows a protection layer such as Cloudflare, Sucuri or Akamai, say it may be a bot block.
4. Content without JavaScript. Read the key pages with site_read_url and text=true. This view is extracted from the HTML the server sends, which is what an AI crawler that does not run JavaScript sees. If the reply says truncated is true, read the page again with max_bytes=1000000, otherwise you will think content near the bottom is missing. For each page, say whether the H1, the first paragraph of the main content, the price or service details and the contact details are in that text, and whether internal_links contains links to the main menu pages.
5. Content with JavaScript. Measure the three most important pages with seo_technical_audit, which renders the page with JavaScript. Flag separately any page where the title, H1 or language differs between the two views. A language difference usually comes from an automatic redirect based on the visitor's country or browser language, and because most AI crawlers come from servers in the United States, it can make them see the wrong language version.
6. What Google sees. If I have a Search Console property, inspect the key pages in one gsc_inspect_urls call: coverage state, last crawl date and Google's chosen canonical. List any key page blocked by robots.txt or shown as "Crawled, currently not indexed". List the submitted sitemaps with their warnings and errors using gsc_sitemaps.
7. Discovery files. Read sitemap.xml and llms.txt with site_read_url. Check that my key pages are in the sitemap and that the lastmod values look real, and say so if every URL carries the same date. If llms.txt is missing, note it as a gap but not a critical one, because none of the major AI search engines has officially said it uses this file as a ranking or citation signal.
State the limits of the tools plainly. site_read_url sends its request as a browser and cannot pose as GPTBot or ClaudeBot, so a firewall rule that blocks only by crawler name will not show up in this audit. If I use a CDN or firewall, remind me to check its panel for a setting that blocks AI crawlers (in Cloudflare, "AI Crawl Control" or "Block AI bots"). If the site is hosted on Opus Growth, do not send me to any CDN panel, and if you see a block, suggest reporting it with report_issue.
Format the output like this. First a three-sentence summary: is my site open to AI search, and if not, what is the biggest blocker. Then a crawler table: crawler, role, robots.txt decision, the rule line that decides it, severity. Then a page table: URL, status code, redirects, x-robots-tag, H1 and first paragraph in the no-JavaScript text (yes or no), difference from the JavaScript view, Google coverage state. End with a fix list ranked by severity, each with its evidence and who should do it (me, a developer, the hosting provider). Where the data is not enough for a conclusion, do not guess, say what is missing.