All crawlers (56)

Grouped by what the crawl is for. Machine copy: agents.json · agents.csv · user-agents.txt

AI training crawlers (13)

Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today.

Crawlerrobots.txt tokenOperatorrobots.txt
anthropic-aianthropic-aiAnthropicn-a
Applebot-ExtendedApplebot-ExtendedApplen-a
BytespiderBytespiderByteDancedisputed
ClaudeBotClaudeBotAnthropicdocumented
cohere-training-data-crawlercohere-training-data-crawlerCoheredocumented
FacebookBotFacebookBotMetadocumented
Google-ExtendedGoogle-ExtendedGooglen-a
GoogleOtherGoogleOtherGoogledocumented
GPTBotGPTBotOpenAIdocumented
meta-externalagentmeta-externalagentMetadocumented
SemrushBot-OCOBSemrushBot-OCOBSemrushdocumented
TikTokSpiderTikTokSpiderByteDancedisputed
Webzio-ExtendedWebzio-ExtendedWebz.iodocumented

Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space.

Crawlerrobots.txt tokenOperatorrobots.txt
AmazonbotAmazonbotAmazondocumented
Claude-SearchBotClaude-SearchBotAnthropicdocumented
Claude-WebClaude-WebAnthropicn-a
DuckAssistBotDuckAssistBotDuckDuckGodocumented
Google-CloudVertexBotGoogle-CloudVertexBotGoogledocumented
OAI-SearchBotOAI-SearchBotOpenAIdocumented
PerplexityBotPerplexityBotPerplexitydocumented
YouBotYouBotYou.comdocumented

User-triggered fetchers (6)

Fetch one page because a person asked for it, right then. One human intent, one request. Blocking them produces a visible error for a real reader.

Crawlerrobots.txt tokenOperatorrobots.txt
ChatGPT-UserChatGPT-UserOpenAIdocumented
Claude-UserClaude-UserAnthropicdocumented
cohere-aicohere-aiCoheredocumented
meta-externalfetchermeta-externalfetcherMetadocumented
MistralAI-UserMistralAI-UserMistral AIdocumented
Perplexity-UserPerplexity-UserPerplexityby-design-no

Corpus and dataset builders (8)

Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect.

Crawlerrobots.txt tokenOperatorrobots.txt
AI2BotAI2BotAllen Institute for AIdocumented
Ai2Bot-DolmaAi2Bot-DolmaAllen Institute for AIdocumented
CCBotCCBotCommon Crawldocumented
DiffbotDiffbotDiffbotdocumented
ImagesiftBotImagesiftBotHive AIdocumented
img2datasetimg2datasetLAION / img2datasetdocumented
omgiliomgiliWebz.iodocumented
omgilibotomgilibotWebz.iodocumented

Classic index-and-rank crawlers. Several also feed their operator's generative answers, which is why the AI opt-out for Google and Apple is a token rather than a block.

Crawlerrobots.txt tokenOperatorrobots.txt
ApplebotApplebotAppledocumented
BaiduspiderBaiduspiderBaidudocumented
bingbotbingbotMicrosoftdocumented
DuckDuckBotDuckDuckBotDuckDuckGodocumented
GooglebotGooglebotGoogledocumented
Googlebot-ImageGooglebot-ImageGoogledocumented
Googlebot-NewsGooglebot-NewsGoogledocumented
PetalBotPetalBotHuaweidocumented
SeznamBotSeznamBotSeznamdocumented
Storebot-GoogleStorebot-GoogleGoogledocumented
TimpibotTimpibotTimpidocumented
YandexBotYandexBotYandexdocumented
YetiYetiNaverdocumented

SEO and backlink crawlers (2)

Commercial link-graph tooling. No user-facing effect either way, and usually a large share of your bot bandwidth.

Crawlerrobots.txt tokenOperatorrobots.txt
AhrefsBotAhrefsBotAhrefsdocumented
SemrushBotSemrushBotSemrushdocumented

Archivers (2)

Preservation crawlers. Their output is public and permanent, which makes them a separate decision from the AI one.

Crawlerrobots.txt tokenOperatorrobots.txt
archive.org_botarchive.org_botInternet Archivedocumented
ia_archiveria_archiverInternet Archivedocumented

Link preview fetchers (1)

Read your Open Graph tags when someone shares a link. Blocking these is almost always an accident.

Crawlerrobots.txt tokenOperatorrobots.txt
facebookexternalhitfacebookexternalhitMetadocumented

Tools and frameworks (3)

Not operators: crawling software anyone can run. The party behind the request is unknown, so treat them as a rate-limit question rather than a consent question.

Crawlerrobots.txt tokenOperatorrobots.txt
FirecrawlAgentFirecrawlAgentFirecrawldocumented
Google-InspectionToolGoogle-InspectionToolGoogledocumented
ScrapyScrapyScrapy projectdocumented