Block every AI crawler

Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed.

curl -s https://www.pathwren.workers.dev/robots/block-all-ai.txt >> robots.txt

The maximal AI opt-out that still leaves you in Google and Bing. Understand the price before deploying it: you will not be cited by any assistant, and when a reader explicitly asks ChatGPT or Claude to open your page, they get an error. Note also that Perplexity-User and Bytespider are listed here but documented as not governed by robots.txt, so this file is a statement of intent for those two, not an enforcement mechanism.

Names 35 crawlers

AI2Bot · Ai2Bot-Dolma · Amazonbot · anthropic-ai · Applebot-Extended · Bytespider · CCBot · ChatGPT-User · Claude-SearchBot · Claude-User · Claude-Web · ClaudeBot · cohere-ai · cohere-training-data-crawler · Diffbot · DuckAssistBot · FacebookBot · Google-CloudVertexBot · Google-Extended · GoogleOther · GPTBot · ImagesiftBot · img2dataset · meta-externalagent · meta-externalfetcher · MistralAI-User · OAI-SearchBot · omgili · omgilibot · Perplexity-User · PerplexityBot · SemrushBot-OCOB · TikTokSpider · Webzio-Extended · YouBot

The file

/robots/block-all-ai.txt · json

# AI Crawler Index — policy: block-all-ai
# Block every AI crawler
# Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed.
# Generated 2026-09-01 from https://www.pathwren.workers.dev/policy/block-all-ai.html
# 35 crawlers named. Paste into robots.txt at your document root.

User-agent: AI2Bot
Disallow: /

User-agent: Ai2Bot-Dolma
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: anthropic-ai   # control token, no crawler uses this user-agent
Disallow: /

User-agent: Applebot-Extended   # control token, no crawler uses this user-agent
Disallow: /

User-agent: Bytespider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: Claude-Web   # control token, no crawler uses this user-agent
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: cohere-ai
Disallow: /

User-agent: cohere-training-data-crawler
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: DuckAssistBot
Disallow: /

User-agent: FacebookBot
Disallow: /

User-agent: Google-CloudVertexBot
Disallow: /

User-agent: Google-Extended   # control token, no crawler uses this user-agent
Disallow: /

User-agent: GoogleOther
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: ImagesiftBot
Disallow: /

User-agent: img2dataset
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: meta-externalfetcher
Disallow: /

User-agent: MistralAI-User
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: omgili
Disallow: /

User-agent: omgilibot
Disallow: /

User-agent: Perplexity-User   # operator states robots.txt does not apply; enforce at the edge
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: SemrushBot-OCOB
Disallow: /

User-agent: TikTokSpider   # compliance disputed; enforce at the edge
Disallow: /

User-agent: Webzio-Extended
Disallow: /

User-agent: YouBot
Disallow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml