lintpage
~/ai-crawlers
§ reference · 34 crawlers · verified 2026-09-16

Every AI crawler, with its source attached.

AI crawler lists copy each other. Strings get invented, versions go stale, and a token that controls nothing gets pasted into ten thousand robots.txt files. Every row here comes from the operator’s own documentation, linked. Where a vendor publishes no user-agent string, this page says so instead of filling the gap.

§ the distinction that matters

Training is a choice. Retrieval is a mistake.

Blocking a training crawler keeps you out of a future model and costs you nothing today. Blocking a live-retrieval crawler makes you invisible in the answer someone is asking for right now. The names are confusingly similar by accident, not design: GPTBot and ChatGPT-User are opposites.

live retrieval · 10
Fetches your page right now because a human asked a question that needs it. Blocking these is what makes you invisible in the answer, and it is almost always an accident.
search index · 11
Builds an index the assistant cites from. Blocking one removes you from that product's answers, but not from the others.
training · 10
Collects content to train future models. Opting out is a legitimate choice that costs you nothing in live AI visibility.
dataset · 2
Collects content into a corpus other people train on. Blocking works going forward only: what is already published stays published.
ads · 1
Checks advertising landing pages. No bearing on organic AI visibility.
§ the table

34 documented crawlers.

Grouped by what blocking one actually costs you. The robots.txt column reports what the operator documents about its own behaviour, not what anyone has observed.

live retrieval

tokenwhat it doesrobots.txtsource
ChatGPT-UserOpenAIFetches a page on demand when a ChatGPT user, Custom GPT or GPT Action asks for it.block it: ChatGPT cannot retrieve your page to answer a question about it. This is the expensive one.may ignorevendor doc →
Claude-UserAnthropicFetches a page when someone asks Claude a question that needs it.block it: Claude cannot retrieve your content in response to a user query.obeysvendor doc →
Perplexity-UserPerplexityFetches a page because a Perplexity user asked something about it.block it: Perplexity cannot pull up your page to answer a live question.may ignorevendor doc →
Google-GeminiNotebookGoogleFetches URLs a Gemini Notebook user supplies as a source.block it: Nothing you can control by robots.txt: as a user-triggered fetcher it is documented to ignore it.may ignorevendor doc →
Google-AgentGoogleAgents on Google infrastructure navigating the web and acting on a user's request.block it: Again not robots-controllable. Google pairs it with a Web Bot Auth experiment under the identity https://agent.bot.goog.may ignorevendor doc →
DuckAssistBotDuckDuckGoCrawls in real time for AI-assisted answers. "This data is not used in any way to train AI models."block it: No AI-assisted answers from your content. "Opting out of DuckAssistBot does not impact organic search rankings." Takes effect after 72 hours.obeysvendor doc →
meta-externalfetcherMetaFetches links a user asked about, including for agentic AI.block it: Meta AI cannot follow a link to your page on a user's behalf.may ignorevendor doc →
facebookexternalhitMetaScrapes a page to build the share preview when someone posts your link.block it: Links to your site post without a title, description or image. Nothing to do with AI.may ignorevendor doc →
Amzn-UserAmazonLive retrieval on a user's behalf, including for Alexa.block it: Alexa cannot read your page out in answer to a question.partlyvendor doc →
MistralAI-UserMistralFetches pages to answer a user's question in Le Chat.block it: Mistral cannot retrieve your page for a live answer.obeysvendor doc →

search index

tokenwhat it doesrobots.txtsource
OAI-SearchBotOpenAIBuilds the search index behind ChatGPT search.block it: You "will not be shown in ChatGPT search answers, though can still appear as navigational links".obeysvendor doc →
Claude-SearchBotAnthropicIndexes content to improve Claude search result quality.block it: You are excluded from the index Claude search draws on.obeysvendor doc →
PerplexityBotPerplexityIndexes pages for Perplexity search. "Not used to crawl content for AI foundation models."block it: You are excluded from Perplexity search results and citations.obeysvendor doc →
GooglebotGoogleThe Search crawler. AI Overviews and AI Mode are built on Search, so this is what feeds them.block it: You leave Google Search entirely. This is the only documented way out of AI Overviews, which is why it is almost never the right move.obeysvendor doc →
ApplebotApplePowers Spotlight, Siri and Safari search. The crawled data also feeds foundation-model training and AI grounding.block it: Removes you from Apple search surfaces.obeysvendor doc →
bingbotMicrosoftBuilds the Bing index. Copilot grounds on indexed content, so this is the Copilot crawler in practice.block it: You leave the Bing index, and with it Copilot citation: Bing's AI Performance report "reflects only content that is eligible for indexing".obeysvendor doc →
meta-webindexerMetaBuilds the index Meta AI search cites from.block it: You are excluded from Meta AI search citations.obeysvendor doc →
Amzn-SearchBotAmazonBuilds Amazon's search index.block it: Excluded from Amazon search results.obeysvendor doc →
MistralAI-IndexMistralBuilds a search index. Explicitly not training.block it: Excluded from Mistral's search index.obeysvendor doc →
YouBotYou.comIndexes pages for You.com search.block it: Excluded from You.com results.obeysvendor doc →
DiffbotDiffbotCrawls for Diffbot's Knowledge Graph. "It is not used for AI training."block it: Excluded from the Knowledge Graph and products built on it.partlyvendor doc →

training

tokenwhat it doesrobots.txtsource
GPTBotOpenAICrawls content for training generative AI foundation models.block it: Your content is excluded from future model training. Live ChatGPT answers are unaffected.obeysvendor doc →
ClaudeBotAnthropicCollects web content that "could potentially contribute to their training".block it: Your content is excluded from Claude training data. Live Claude answers are unaffected.obeysvendor doc →
anthropic-aiAnthropicLegacy token in wide circulation. Anthropic does not document it.block it: Unknown. Anthropic lists only ClaudeBot, Claude-User and Claude-SearchBot, so this token may control nothing.unknownnot documented
Google-ExtendedGoogleControls training of future Gemini models and grounding in Gemini Apps and Vertex AI.block it: Excluded from Gemini training and grounding. It does NOT remove you from AI Overviews, and it is not a Search ranking signal.obeysvendor doc →
Google-CloudVertexBotGoogleCrawls sites on their owner's request to build Vertex AI Agents.block it: Vertex AI Agents cannot ingest your site. "It has no effect on Google Search or other products."obeysvendor doc →
GoogleOtherGoogleGeneric crawler used for one-off internal research and development.block it: Preferences for GoogleOther "don't affect any specific product", so blocking it changes little you can observe.obeysvendor doc →
Applebot-ExtendedAppleOpts your content out of training Apple's general-purpose foundation models.block it: Excluded from Apple Intelligence training only. "Webpages that disallow Applebot-Extended can still be included in search results", and it is not considered in Search ranking.obeysvendor doc →
meta-externalagentMetaTrains foundation AI models, and indexes content directly to improve products.block it: Your content is excluded from Meta AI training and indexing.obeysvendor doc →
AmazonbotAmazonCrawls to "improve our products and services", which may include training Amazon AI.block it: Excluded from Amazon AI training and product data.obeysvendor doc →
MistralAI-TrainingMistralCollects training data for Mistral models.block it: Excluded from Mistral training data.obeysvendor doc →

dataset

tokenwhat it doesrobots.txtsource
CCBotCommon CrawlBuilds the Common Crawl public archive. A nonprofit corpus, not an AI product.block it: You leave the corpus that many third-party models train on, which blocks them indirectly. Archives already published stay published.obeysvendor doc →
Webzio-extendedWebz.ioCollects content resold as datasets. The "-extended" token covers AI and ML use.block it: Your content is kept out of datasets sold on for AI training.obeysvendor doc →

ads

tokenwhat it doesrobots.txtsource
OAI-AdsBotOpenAIValidates advertising landing pages. "Only visits pages submitted as ads"; the data is not used for training.block it: Ad landing pages you submit cannot be validated. No effect on organic AI visibility.obeysvendor doc →
§ the strings

User-agent headers, as published.

Quoted verbatim, including the typos. Perplexity’s strings have an unbalanced parenthesis and Meta’s carry a relative URL; both are reproduced as the vendor prints them, because a matcher built on a tidied-up string can miss. Match the token, not the whole header.

GPTBotexample only, version moves
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

robots.txt: Disallowing GPTBot "indicates a site's content should not be used in training".

OAI-SearchBotexample only, version moves
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot

When fetching robots.txt itself OpenAI inserts a "robots.txt;" marker into the string.

ChatGPT-Userdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot

The self-link in the string is +https://openai.com/bot, not /chatgpt-user.

robots.txt: "Because these actions are initiated by a user, robots.txt rules may not apply."

OAI-AdsBotdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot
ClaudeBotoperator publishes none
The operator publishes no user-agent string. Any string you have seen for this crawler came from a third party.

Anthropic publishes no User-Agent header string for any of its bots. Match the token.

robots.txt: Honours robots.txt and the non-standard Crawl-delay directive.

Claude-Useroperator publishes none
The operator publishes no user-agent string. Any string you have seen for this crawler came from a third party.

robots.txt: No user-initiated exception is documented, unlike ChatGPT-User and Perplexity-User.

Claude-SearchBotoperator publishes none
The operator publishes no user-agent string. Any string you have seen for this crawler came from a third party.
anthropic-aioperator publishes none
The operator publishes no user-agent string. Any string you have seen for this crawler came from a third party.

Appears in countless robots.txt files and blocklists. Absent from Anthropic's documentation, with no retirement notice.

PerplexityBotdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Perplexity-Userdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

robots.txt: "Since a user requested the fetch, this fetcher generally ignores robots.txt rules."

Google-Extendedrobots.txt token, never sent as a header
No request header exists. This token is a control you write in robots.txt; nothing ever arrives with it as a User-Agent.

A control token, not a fetcher. Nothing ever arrives with this User-Agent, so the checker reads your robots.txt for it and sends no request: there is no live result to report.

robots.txt: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."

Googlebotdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36

Listed as the contrast case: the AI control people reach for is Google-Extended, but the AI surface they mean is governed here.

robots.txt: Common crawlers "always respect robots.txt rules for automatic crawls".

Google-CloudVertexBotoperator publishes none
The operator publishes no user-agent string. Any string you have seen for this crawler came from a third party.

Google documents the substring only, not a full header string.

GoogleOtherdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GoogleOther) Chrome/W.X.Y.Z Safari/537.36
Google-GeminiNotebookdocumented
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)

Replaced Google-NotebookLM, which Google supported until August 2026.

robots.txt: User-triggered fetchers "generally ignore robots.txt rules".

Google-Agentdocumented
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko; compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent) Chrome/W.X.Y.Z Safari/537.36

robots.txt: User-triggered fetchers "generally ignore robots.txt rules".

Applebot-Extendedrobots.txt token, never sent as a header
No request header exists. This token is a control you write in robots.txt; nothing ever arrives with it as a User-Agent.

Training only. Grounding is a separate control: nosnippet stops Apple using the page as AI context.

Applebotexample only, version moves
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)

robots.txt: Respects robots.txt, falls back to Googlebot rules when Applebot is not named, and does not follow crawl-delay.

bingbotdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36

Microsoft controls AI use by meta tag rather than by token: noindex bars training, noarchive bars Chat and Copilot, nocache limits them to URL, title and snippet.

robots.txt: Honours one group in priority order: bingbot, then msnbot, then the wildcard.

DuckAssistBotdocumented
DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)
meta-externalagentdocumented
meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)

The parenthetical is relativized on Meta's own page; match the "meta-externalagent/1.1" prefix rather than the whole string.

meta-externalfetcherdocumented
meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers)

robots.txt: "May bypass robots.txt because it performs fetches that were requested by the user."

meta-webindexerdocumented
meta-webindexer/1.1 (+/documentation/sharing/webmasters/web-crawlers)
facebookexternalhitdocumented
facebookexternalhit/1.1 (+http://www.facebook.com/externalhit_uatext.php)

Listed because it is routinely swept up in AI blocklists, which silently breaks link previews.

robots.txt: "Might bypass robots.txt when performing security or integrity checks."

Amazonbotdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36

robots.txt: "Respects the Robots Exclusion Protocol."

Amzn-SearchBotdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36
Amzn-Userdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36

robots.txt: "May not follow all robots.txt directives."

MistralAI-Userdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)
MistralAI-Indexdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)
MistralAI-Trainingdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)

The only Mistral agent with no published IP list, so it is the one you cannot verify.

CCBotdocumented
CCBot/2.0 (https://commoncrawl.org/faq/)

Predates the LLM era. Verifiable by reverse DNS under *.crawl.commoncrawl.org.

YouBotdocumented
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; YouBot/1.0; +https://docs.you.com/youbot; env:prod) Chrome/X.X.X.X Safari/537.36

Verifiable three ways: IP range, reverse DNS, and Cloudflare Web Bot Auth.

robots.txt: "Fully respects robots.txt directives, including user-agent specific rules and crawl-delay."

Diffbotoperator publishes none
The operator publishes no user-agent string. Any string you have seen for this crawler came from a third party.

Diffbot documents the token and a Diffbot-User variant, but no header string.

robots.txt: Obeys robots.txt by default, but the default is overridable under a partnership agreement.

Webzio-extendedoperator publishes none
The operator publishes no user-agent string. Any string you have seen for this crawler came from a third party.

Webz.io also operates the Webzio and Omgilibot tokens. No header string is published for any of them.

robots.txt: "Meticulously adheres to robots.txt exclusions."

§ the one everyone gets wrong

Google-Extended does not control AI Overviews.

The standard advice is to disallow Google-Extended to stay out of AI Overviews. It does not work. Google documents that token as covering two things: training future Gemini models, and grounding answers in Gemini Apps and Vertex AI. AI Overviews are not on the list.

AI Overviews are built on Search, so they follow Googlebot, and Google states there are no additional requirements to appear in them. The only documented way out is to leave Google Search. That is the real trade, and it is why the advice is worth correcting rather than repeating: sites disallow Google-Extended, lose their Gemini visibility, stay in AI Overviews, and change nothing about their rankings.

Apple splits the same way. Applebot-Extended opts you out of training only, and Apple says pages that disallow it can still appear in search results. Grounding is a separate control (nosnippet). We shipped the wrong version of the Google claim ourselves until 2026-09-16, which is a good argument for citing the vendor page on every row.

§ unsourced

5 names everyone lists that nobody documents.

These appear in most AI crawler lists, usually with a confident-looking user-agent string. We went to the operator and could not find one. That does not prove the crawler is not real; it means nobody has published what you would need to block it correctly.

?BytespiderByteDance · checked 2026-09-16Listed everywhere as ByteDance's AI training crawler, usually with a full UA string and a warning that it ignores robots.txt.ByteDance publishes no crawler documentation. Its webmaster platform returns an empty JavaScript shell and every help path errors. The circulating UA string has no primary source, and we found no evidence for the robots.txt claim in Cloudflare's published research either - only that Bytespider was the most-blocked AI crawler, with its share falling from 14.1% to 2.4%.
?xAI-Bot / GrokBotxAI · checked 2026-09-16Named in blocklists as the crawler behind Grok.xAI documents no crawler under any name. There is no page on x.ai or docs.x.ai, and the names circulate only through aggregators. xAI does publish Content-Signal preferences in its own robots.txt, which makes it a signal publisher rather than a documented operator.
?cohere-aiCohere · checked 2026-09-16Widely copied into robots.txt as Cohere's AI crawler token.Cohere states plainly: "We do not use Cohere bots or user agents for the purpose of crawling or scraping web content to train generative AI foundation models at this time." The string "cohere-ai" appears on their page only as a GitHub organisation name. Blocking this token blocks nothing Cohere operates.
?Copilot crawlerMicrosoft · checked 2026-09-16Blocklists routinely carry a Copilot-specific token, by analogy with GPTBot and ClaudeBot.Microsoft documents no AI-specific crawler. Its crawler list names only Bingbot, AdIdxBot, BingPreview, MicrosoftPreview and BingVideoPreview, so there is no user-agent to block. Copilot use is controlled by meta tag instead: noindex bars training, noarchive bars Chat and Copilot, nocache limits them to URL, title and snippet.
?FacebookBotMeta · checked 2026-09-16Appears in older blocklists as Meta's AI or speech-recognition crawler.Meta no longer documents it. The page that described it now serves FacebookExternalHit content instead, with no retirement notice. The current Meta crawlers are the meta-* tokens above.
§ verification

A user-agent header is not an identity.

Anything can send any header, so blocking by user-agent is a courtesy system. 15 of the crawlers here publish an IP range list you can check a request against. Reverse DNS and Web Bot Auth are documented for a few more, and where that is the case it is noted on the row itself rather than counted here.

§ faq

Questions, answered.

What is the full list of AI crawler user agents?
The table on this page lists 34 crawlers that their operators document, covering OpenAI, Anthropic, Google, Apple, Microsoft, Perplexity, Meta, Amazon, Mistral, DuckDuckGo, You.com, Common Crawl and others. Each row links to the vendor page it came from. 5 more names in common circulation are listed separately because no operator documents them at all.
Does blocking Google-Extended keep me out of AI Overviews?
No, and this is the most common mistake in AI crawler advice. Google documents Google-Extended as controlling training of future Gemini models and grounding in Gemini Apps and Vertex AI. AI Overviews are part of Search and follow Googlebot, and Google states there are no additional requirements to appear in them. The only documented way out of AI Overviews is to leave Google Search, which almost nobody wants.
Which AI crawlers ignore robots.txt?
Only the ones whose operators say so. ChatGPT-User, Perplexity-User, Meta-ExternalFetcher, FacebookExternalHit and Google's user-triggered fetchers are all documented as possibly ignoring robots.txt, on the reasoning that a human asked for the fetch. Amzn-User may not follow all directives. Every other documented crawler on this page states that it obeys robots.txt. Claims that a particular crawler secretly ignores it are common and usually unsourced.
Is a user-agent string proof of who is crawling me?
No. Any client can send any User-Agent header, so the string alone proves nothing. 15 of the crawlers here publish an IP range list you can check a request against, and each one is linked below. A few others document reverse DNS or Web Bot Auth instead of a list, noted on their own rows. For the rest, a matching header is a claim, not an identity.
How do I check whether AI crawlers can actually reach my site?
Reading your robots.txt is not enough. The common failure in 2026 is a CDN or WAF returning 403 to AI bots at the edge, before your origin sees the request, while your robots.txt says they are welcome. The only way to know is to fetch your page as each bot and read the response. The LintPage AI Crawler Checker does that live.

LintPage is not affiliated with any operator named on this page. Tokens and user-agent strings are quoted from vendor documentation so you can check your own configuration against them. Every row was last verified on 2026-09-16; these change, so follow the source link before relying on one.

§ the measurable part

Can they actually reach you?

A correct robots.txt proves nothing if your CDN returns 403 at the edge. LintPage fetches your page as each bot and reports what it really got back. Free, no signup.

run the AI crawler check →