lintpage
~/tools/ai-crawler-checker

AI Crawler Checker

Most "AI crawler" tools only parse your robots.txt. We actually fetch your site as each bot and tell you what they really see. Catches Cloudflare and WAF blocks that other tools miss.

§ what this tool checks

Rules applied to every scan.

Edge platforms now block AI bots by default in ways your robots.txt cannot see. On 15 September 2026 Cloudflare set new defaults that block AI training and agent traffic on pages showing ads, and applied them to free-plan zones whose owners never changed the setting. A permissive robots.txt says nothing about what the bot actually received. Robots.txt parsers cannot catch this; live-fetch testing can. This tool fetches your page as the 8 AI crawlers that actually send requests and reports what each one really received, and checks the two robots.txt control tokens (Google-Extended, Applebot-Extended) against your file, because nothing ever arrives with those as a User-Agent.

GPTBot, ChatGPT-User, OAI-SearchBot (OpenAI)
ClaudeBot, Claude-User, anthropic-ai (Anthropic)
PerplexityBot, Perplexity-User (Perplexity)
Google-Extended (Gemini training and grounding, not AI Overviews)
Applebot-Extended (Apple Intelligence)
Silent block detection: robots allows but live fetch returns 403/429
JS-rendering heuristic: 200 OK with empty body text
robots.txt parse with per-bot rules and llms.txt check
§ faq

Questions, answered.

Why do robots.txt parsers miss most AI bot blocks today?
Because the block usually happens one layer above robots.txt. Your file can say "User-agent: GPTBot" and "Allow: /" while Cloudflare, a WAF or a CDN rule returns 403 before the request ever reaches your origin. Cloudflare set new AI bot defaults on 15 September 2026 and applied them to free-plan zones that had never changed the setting, so this can be true of a site nobody touched. Parsers only read the file; they never test what the bot receives. This tool fetches your page with each bot's User-Agent and reports the real HTTP response.
What is the difference between GPTBot and ChatGPT-User?
GPTBot is OpenAI's training crawler — it indexes your content for future model training. Opting out is a legitimate choice. ChatGPT-User is the live browsing agent that fetches pages when a ChatGPT user asks a question. Blocking ChatGPT-User removes you from real-time AI answers — almost always unintentional. Same distinction applies to Claude/Claude-User and Perplexity/Perplexity-User.
Is opting out of Google-Extended the same as blocking Googlebot?
No. Google-Extended is a separate token that controls whether Google can use your content to train future Gemini models and to ground answers in Gemini Apps and Vertex AI. Disallowing it has no effect on regular Googlebot or your search rankings. It also does not remove you from AI Overviews, despite the common advice: Google documents AI Overviews as part of Search, governed by Googlebot, with no additional requirements to appear. Leaving Search is the only documented way out of them.
My page returned 200 but the tool says "empty content" — what happened?
Your page returned HTML successfully but contained less than 200 characters of readable text. That usually means the content is hydrated client-side (React, Vue, and so on), so the initial payload is a shell. Whether a given AI crawler can recover from that is genuinely unknown: Googlebot, bingbot and Applebot document that they render JavaScript, while OpenAI, Anthropic, Perplexity and Meta document nothing either way. Server-render the primary content and the question stops mattering.
§ run the full audit

Stop guessing. Scan everything in one click.

60 automated checks across meta tags, robots.txt, Open Graph, sitemaps, headings, AI visibility, and more — free, no signup.

run a full scan →