Every AI crawler, with its source attached.
AI crawler lists copy each other. Strings get invented, versions go stale, and a token that controls nothing gets pasted into ten thousand robots.txt files. Every row here comes from the operator’s own documentation, linked. Where a vendor publishes no user-agent string, this page says so instead of filling the gap.
Training is a choice. Retrieval is a mistake.
Blocking a training crawler keeps you out of a future model and costs you nothing today. Blocking a live-retrieval crawler makes you invisible in the answer someone is asking for right now. The names are confusingly similar by accident, not design: GPTBot and ChatGPT-User are opposites.
- live retrieval · 10
- Fetches your page right now because a human asked a question that needs it. Blocking these is what makes you invisible in the answer, and it is almost always an accident.
- search index · 11
- Builds an index the assistant cites from. Blocking one removes you from that product's answers, but not from the others.
- training · 10
- Collects content to train future models. Opting out is a legitimate choice that costs you nothing in live AI visibility.
- dataset · 2
- Collects content into a corpus other people train on. Blocking works going forward only: what is already published stays published.
- ads · 1
- Checks advertising landing pages. No bearing on organic AI visibility.
34 documented crawlers.
Grouped by what blocking one actually costs you. The robots.txt column reports what the operator documents about its own behaviour, not what anyone has observed.
live retrieval
| token | what it does | robots.txt | source |
|---|---|---|---|
| ChatGPT-UserOpenAI | Fetches a page on demand when a ChatGPT user, Custom GPT or GPT Action asks for it.block it: ChatGPT cannot retrieve your page to answer a question about it. This is the expensive one. | may ignore | vendor doc → |
| Claude-UserAnthropic | Fetches a page when someone asks Claude a question that needs it.block it: Claude cannot retrieve your content in response to a user query. | obeys | vendor doc → |
| Perplexity-UserPerplexity | Fetches a page because a Perplexity user asked something about it.block it: Perplexity cannot pull up your page to answer a live question. | may ignore | vendor doc → |
| Google-GeminiNotebookGoogle | Fetches URLs a Gemini Notebook user supplies as a source.block it: Nothing you can control by robots.txt: as a user-triggered fetcher it is documented to ignore it. | may ignore | vendor doc → |
| Google-AgentGoogle | Agents on Google infrastructure navigating the web and acting on a user's request.block it: Again not robots-controllable. Google pairs it with a Web Bot Auth experiment under the identity https://agent.bot.goog. | may ignore | vendor doc → |
| DuckAssistBotDuckDuckGo | Crawls in real time for AI-assisted answers. "This data is not used in any way to train AI models."block it: No AI-assisted answers from your content. "Opting out of DuckAssistBot does not impact organic search rankings." Takes effect after 72 hours. | obeys | vendor doc → |
| meta-externalfetcherMeta | Fetches links a user asked about, including for agentic AI.block it: Meta AI cannot follow a link to your page on a user's behalf. | may ignore | vendor doc → |
| facebookexternalhitMeta | Scrapes a page to build the share preview when someone posts your link.block it: Links to your site post without a title, description or image. Nothing to do with AI. | may ignore | vendor doc → |
| Amzn-UserAmazon | Live retrieval on a user's behalf, including for Alexa.block it: Alexa cannot read your page out in answer to a question. | partly | vendor doc → |
| MistralAI-UserMistral | Fetches pages to answer a user's question in Le Chat.block it: Mistral cannot retrieve your page for a live answer. | obeys | vendor doc → |
search index
| token | what it does | robots.txt | source |
|---|---|---|---|
| OAI-SearchBotOpenAI | Builds the search index behind ChatGPT search.block it: You "will not be shown in ChatGPT search answers, though can still appear as navigational links". | obeys | vendor doc → |
| Claude-SearchBotAnthropic | Indexes content to improve Claude search result quality.block it: You are excluded from the index Claude search draws on. | obeys | vendor doc → |
| PerplexityBotPerplexity | Indexes pages for Perplexity search. "Not used to crawl content for AI foundation models."block it: You are excluded from Perplexity search results and citations. | obeys | vendor doc → |
| GooglebotGoogle | The Search crawler. AI Overviews and AI Mode are built on Search, so this is what feeds them.block it: You leave Google Search entirely. This is the only documented way out of AI Overviews, which is why it is almost never the right move. | obeys | vendor doc → |
| ApplebotApple | Powers Spotlight, Siri and Safari search. The crawled data also feeds foundation-model training and AI grounding.block it: Removes you from Apple search surfaces. | obeys | vendor doc → |
| bingbotMicrosoft | Builds the Bing index. Copilot grounds on indexed content, so this is the Copilot crawler in practice.block it: You leave the Bing index, and with it Copilot citation: Bing's AI Performance report "reflects only content that is eligible for indexing". | obeys | vendor doc → |
| meta-webindexerMeta | Builds the index Meta AI search cites from.block it: You are excluded from Meta AI search citations. | obeys | vendor doc → |
| Amzn-SearchBotAmazon | Builds Amazon's search index.block it: Excluded from Amazon search results. | obeys | vendor doc → |
| MistralAI-IndexMistral | Builds a search index. Explicitly not training.block it: Excluded from Mistral's search index. | obeys | vendor doc → |
| YouBotYou.com | Indexes pages for You.com search.block it: Excluded from You.com results. | obeys | vendor doc → |
| DiffbotDiffbot | Crawls for Diffbot's Knowledge Graph. "It is not used for AI training."block it: Excluded from the Knowledge Graph and products built on it. | partly | vendor doc → |
training
| token | what it does | robots.txt | source |
|---|---|---|---|
| GPTBotOpenAI | Crawls content for training generative AI foundation models.block it: Your content is excluded from future model training. Live ChatGPT answers are unaffected. | obeys | vendor doc → |
| ClaudeBotAnthropic | Collects web content that "could potentially contribute to their training".block it: Your content is excluded from Claude training data. Live Claude answers are unaffected. | obeys | vendor doc → |
| anthropic-aiAnthropic | Legacy token in wide circulation. Anthropic does not document it.block it: Unknown. Anthropic lists only ClaudeBot, Claude-User and Claude-SearchBot, so this token may control nothing. | unknown | not documented |
| Google-ExtendedGoogle | Controls training of future Gemini models and grounding in Gemini Apps and Vertex AI.block it: Excluded from Gemini training and grounding. It does NOT remove you from AI Overviews, and it is not a Search ranking signal. | obeys | vendor doc → |
| Google-CloudVertexBotGoogle | Crawls sites on their owner's request to build Vertex AI Agents.block it: Vertex AI Agents cannot ingest your site. "It has no effect on Google Search or other products." | obeys | vendor doc → |
| GoogleOtherGoogle | Generic crawler used for one-off internal research and development.block it: Preferences for GoogleOther "don't affect any specific product", so blocking it changes little you can observe. | obeys | vendor doc → |
| Applebot-ExtendedApple | Opts your content out of training Apple's general-purpose foundation models.block it: Excluded from Apple Intelligence training only. "Webpages that disallow Applebot-Extended can still be included in search results", and it is not considered in Search ranking. | obeys | vendor doc → |
| meta-externalagentMeta | Trains foundation AI models, and indexes content directly to improve products.block it: Your content is excluded from Meta AI training and indexing. | obeys | vendor doc → |
| AmazonbotAmazon | Crawls to "improve our products and services", which may include training Amazon AI.block it: Excluded from Amazon AI training and product data. | obeys | vendor doc → |
| MistralAI-TrainingMistral | Collects training data for Mistral models.block it: Excluded from Mistral training data. | obeys | vendor doc → |
dataset
| token | what it does | robots.txt | source |
|---|---|---|---|
| CCBotCommon Crawl | Builds the Common Crawl public archive. A nonprofit corpus, not an AI product.block it: You leave the corpus that many third-party models train on, which blocks them indirectly. Archives already published stay published. | obeys | vendor doc → |
| Webzio-extendedWebz.io | Collects content resold as datasets. The "-extended" token covers AI and ML use.block it: Your content is kept out of datasets sold on for AI training. | obeys | vendor doc → |
ads
| token | what it does | robots.txt | source |
|---|---|---|---|
| OAI-AdsBotOpenAI | Validates advertising landing pages. "Only visits pages submitted as ads"; the data is not used for training.block it: Ad landing pages you submit cannot be validated. No effect on organic AI visibility. | obeys | vendor doc → |
User-agent headers, as published.
Quoted verbatim, including the typos. Perplexity’s strings have an unbalanced parenthesis and Meta’s carry a relative URL; both are reproduced as the vendor prints them, because a matcher built on a tidied-up string can miss. Match the token, not the whole header.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbotrobots.txt: Disallowing GPTBot "indicates a site's content should not be used in training".
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbotWhen fetching robots.txt itself OpenAI inserts a "robots.txt;" marker into the string.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/botThe self-link in the string is +https://openai.com/bot, not /chatgpt-user.
robots.txt: "Because these actions are initiated by a user, robots.txt rules may not apply."
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbotAnthropic publishes no User-Agent header string for any of its bots. Match the token.
robots.txt: Honours robots.txt and the non-standard Crawl-delay directive.
robots.txt: No user-initiated exception is documented, unlike ChatGPT-User and Perplexity-User.
Appears in countless robots.txt files and blocklists. Absent from Anthropic's documentation, with no retirement notice.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)robots.txt: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules."
A control token, not a fetcher. Nothing ever arrives with this User-Agent, so the checker reads your robots.txt for it and sends no request: there is no live result to report.
robots.txt: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36Listed as the contrast case: the AI control people reach for is Google-Extended, but the AI surface they mean is governed here.
robots.txt: Common crawlers "always respect robots.txt rules for automatic crawls".
Google documents the substring only, not a full header string.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GoogleOther) Chrome/W.X.Y.Z Safari/537.36Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)Replaced Google-NotebookLM, which Google supported until August 2026.
robots.txt: User-triggered fetchers "generally ignore robots.txt rules".
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko; compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent) Chrome/W.X.Y.Z Safari/537.36robots.txt: User-triggered fetchers "generally ignore robots.txt rules".
Training only. Grounding is a separate control: nosnippet stops Apple using the page as AI context.
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)robots.txt: Respects robots.txt, falls back to Googlebot rules when Applebot is not named, and does not follow crawl-delay.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36Microsoft controls AI use by meta tag rather than by token: noindex bars training, noarchive bars Chat and Copilot, nocache limits them to URL, title and snippet.
robots.txt: Honours one group in priority order: bingbot, then msnbot, then the wildcard.
DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)The parenthetical is relativized on Meta's own page; match the "meta-externalagent/1.1" prefix rather than the whole string.
meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers)robots.txt: "May bypass robots.txt because it performs fetches that were requested by the user."
meta-webindexer/1.1 (+/documentation/sharing/webmasters/web-crawlers)facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)Listed because it is routinely swept up in AI blocklists, which silently breaks link previews.
robots.txt: "Might bypass robots.txt when performing security or integrity checks."
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36robots.txt: "Respects the Robots Exclusion Protocol."
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36robots.txt: "May not follow all robots.txt directives."
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)The only Mistral agent with no published IP list, so it is the one you cannot verify.
CCBot/2.0 (https://commoncrawl.org/faq/)Predates the LLM era. Verifiable by reverse DNS under *.crawl.commoncrawl.org.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; YouBot/1.0; +https://docs.you.com/youbot; env:prod) Chrome/X.X.X.X Safari/537.36Verifiable three ways: IP range, reverse DNS, and Cloudflare Web Bot Auth.
robots.txt: "Fully respects robots.txt directives, including user-agent specific rules and crawl-delay."
Diffbot documents the token and a Diffbot-User variant, but no header string.
robots.txt: Obeys robots.txt by default, but the default is overridable under a partnership agreement.
Webz.io also operates the Webzio and Omgilibot tokens. No header string is published for any of them.
robots.txt: "Meticulously adheres to robots.txt exclusions."
Google-Extended does not control AI Overviews.
The standard advice is to disallow Google-Extended to stay out of AI Overviews. It does not work. Google documents that token as covering two things: training future Gemini models, and grounding answers in Gemini Apps and Vertex AI. AI Overviews are not on the list.
AI Overviews are built on Search, so they follow Googlebot, and Google states there are no additional requirements to appear in them. The only documented way out is to leave Google Search. That is the real trade, and it is why the advice is worth correcting rather than repeating: sites disallow Google-Extended, lose their Gemini visibility, stay in AI Overviews, and change nothing about their rankings.
Apple splits the same way. Applebot-Extended opts you out of training only, and Apple says pages that disallow it can still appear in search results. Grounding is a separate control (nosnippet). We shipped the wrong version of the Google claim ourselves until 2026-09-16, which is a good argument for citing the vendor page on every row.
5 names everyone lists that nobody documents.
These appear in most AI crawler lists, usually with a confident-looking user-agent string. We went to the operator and could not find one. That does not prove the crawler is not real; it means nobody has published what you would need to block it correctly.
A user-agent header is not an identity.
Anything can send any header, so blocking by user-agent is a courtesy system. 15 of the crawlers here publish an IP range list you can check a request against. Reverse DNS and Web Bot Auth are documented for a few more, and where that is the case it is noted on the row itself rather than counted here.
- › GPTBot published IP list →
- › OAI-SearchBot published IP list →
- › ChatGPT-User published IP list →
- › OAI-AdsBot published IP list →
- › ClaudeBot published IP list →
- › Claude-User published IP list →
- › Claude-SearchBot published IP list →
- › PerplexityBot published IP list →
- › Perplexity-User published IP list →
- › Amazonbot published IP list →
- › Amzn-SearchBot published IP list →
- › Amzn-User published IP list →
- › MistralAI-User published IP list →
- › MistralAI-Index published IP list →
- › CCBot published IP list →
What people actually need to decide.
A list of user agents does not tell you whether to block one, or what a block is worth. These do, from the same vendor documentation.
Questions, answered.
What is the full list of AI crawler user agents?
Does blocking Google-Extended keep me out of AI Overviews?
Which AI crawlers ignore robots.txt?
Is a user-agent string proof of who is crawling me?
How do I check whether AI crawlers can actually reach my site?
LintPage is not affiliated with any operator named on this page. Tokens and user-agent strings are quoted from vendor documentation so you can check your own configuration against them. Every row was last verified on 2026-09-16; these change, so follow the source link before relying on one.
Can they actually reach you?
A correct robots.txt proves nothing if your CDN returns 403 at the edge. LintPage fetches your page as each bot and reports what it really got back. Free, no signup.
run the AI crawler check →