Allow everything, explicitly

Every crawler on this index is named and allowed. Use when you want maximum reach into search and assistants and have nothing to withhold.

curl -s https://www.pathwren.workers.dev/robots/allow-all.txt >> robots.txt

An empty robots.txt already allows everything, so this file is not about permission — it is about being explicit. Naming each token means a later change is a one-line diff instead of a rewrite, and it documents that the allow was a decision. This is the policy this site itself serves.

Names 56 crawlers

AhrefsBot · AI2Bot · Ai2Bot-Dolma · Amazonbot · anthropic-ai · Applebot · Applebot-Extended · archive.org_bot · Baiduspider · bingbot · Bytespider · CCBot · ChatGPT-User · Claude-SearchBot · Claude-User · Claude-Web · ClaudeBot · cohere-ai · cohere-training-data-crawler · Diffbot · DuckAssistBot · DuckDuckBot · FacebookBot · facebookexternalhit · FirecrawlAgent · Google-CloudVertexBot · Google-Extended · Google-InspectionTool · Googlebot · Googlebot-Image · Googlebot-News · GoogleOther · GPTBot · ia_archiver · ImagesiftBot · img2dataset · meta-externalagent · meta-externalfetcher · MistralAI-User · OAI-SearchBot · omgili · omgilibot · Perplexity-User · PerplexityBot · PetalBot · Scrapy · SemrushBot · SemrushBot-OCOB · SeznamBot · Storebot-Google · TikTokSpider · Timpibot · Webzio-Extended · YandexBot · Yeti · YouBot

The file

/robots/allow-all.txt · json

# AI Crawler Index — policy: allow-all
# Allow everything, explicitly
# Every crawler on this index is named and allowed. Use when you want maximum reach into search and assistants and have nothing to withhold.
# Generated 2026-09-01 from https://www.pathwren.workers.dev/policy/allow-all.html
# 56 crawlers named. Paste into robots.txt at your document root.

User-agent: AhrefsBot
Allow: /

User-agent: AI2Bot
Allow: /

User-agent: Ai2Bot-Dolma
Allow: /

User-agent: Amazonbot
Allow: /

User-agent: anthropic-ai   # control token, no crawler uses this user-agent
Allow: /

User-agent: Applebot
Allow: /

User-agent: Applebot-Extended   # control token, no crawler uses this user-agent
Allow: /

User-agent: archive.org_bot
Allow: /

User-agent: Baiduspider
Allow: /

User-agent: bingbot
Allow: /

User-agent: Bytespider   # compliance disputed; enforce at the edge
Allow: /

User-agent: CCBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Claude-Web   # control token, no crawler uses this user-agent
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: cohere-ai
Allow: /

User-agent: cohere-training-data-crawler
Allow: /

User-agent: Diffbot
Allow: /

User-agent: DuckAssistBot
Allow: /

User-agent: DuckDuckBot
Allow: /

User-agent: FacebookBot
Allow: /

User-agent: facebookexternalhit
Allow: /

User-agent: FirecrawlAgent
Allow: /

User-agent: Google-CloudVertexBot
Allow: /

User-agent: Google-Extended   # control token, no crawler uses this user-agent
Allow: /

User-agent: Google-InspectionTool
Allow: /

User-agent: Googlebot
Allow: /

User-agent: Googlebot-Image
Allow: /

User-agent: Googlebot-News
Allow: /

User-agent: GoogleOther
Allow: /

User-agent: GPTBot
Allow: /

User-agent: ia_archiver
Allow: /

User-agent: ImagesiftBot
Allow: /

User-agent: img2dataset
Allow: /

User-agent: meta-externalagent
Allow: /

User-agent: meta-externalfetcher
Allow: /

User-agent: MistralAI-User
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: omgili
Allow: /

User-agent: omgilibot
Allow: /

User-agent: Perplexity-User   # operator states robots.txt does not apply; enforce at the edge
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: PetalBot
Allow: /

User-agent: Scrapy
Allow: /

User-agent: SemrushBot
Allow: /

User-agent: SemrushBot-OCOB
Allow: /

User-agent: SeznamBot
Allow: /

User-agent: Storebot-Google
Allow: /

User-agent: TikTokSpider   # compliance disputed; enforce at the edge
Allow: /

User-agent: Timpibot
Allow: /

User-agent: Webzio-Extended
Allow: /

User-agent: YandexBot
Allow: /

User-agent: Yeti
Allow: /

User-agent: YouBot
Allow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml