AI training crawlers

Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today.

CrawlerTokenOperatorCost of blocking
anthropic-aianthropic-aiAnthropicNone. Nothing crawls under this name today; keeping the rule is harmless insurance.…
Applebot-ExtendedApplebot-ExtendedAppleExcluded from Apple Intelligence training. Siri, Spotlight and Safari suggestions are unaf…
BytespiderBytespiderByteDanceLittle to lose. If you want it gone, expect to block by user-agent at the edge rather than…
ClaudeBotClaudeBotAnthropicContent excluded from training data for future Claude models. No effect on Claude's abilit…
cohere-training-data-crawlercohere-training-data-crawlerCohereExcluded from Cohere model training.…
FacebookBotFacebookBotMetaNegligible today. Keep the rule; expect little traffic.…
Google-ExtendedGoogle-ExtendedGoogleYou are excluded from Gemini grounding and Gemini training. Google Search ranking and inde…
GoogleOtherGoogleOtherGoogleNo effect on Search indexing. Blocks internal Google research and product fetches.…
GPTBotGPTBotOpenAIYour content is excluded from training data for future OpenAI models. No effect on ChatGPT…
meta-externalagentmeta-externalagentMetaExcluded from Meta AI training. Link previews on Facebook, Instagram and WhatsApp are unaf…
SemrushBot-OCOBSemrushBot-OCOBSemrushExclusion from Semrush's AI corpus, with its SEO crawl unaffected.…
TikTokSpiderTikTokSpiderByteDanceLittle to lose unless TikTok search referral matters to you.…
Webzio-ExtendedWebzio-ExtendedWebz.ioYour content is excluded from the AI-training tier of Webz.io's product while ordinary col…

json