Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot.
curl -s https://www.pathwren.workers.dev/robots/block-datasets.txt >> robots.txt
These are the highest-leverage blocks per line, because one crawl becomes many downstream training runs. It is also the block with the longest delay before it has any effect, and no effect at all on archives already published.
AI2Bot · Ai2Bot-Dolma · CCBot · Diffbot · ImagesiftBot · img2dataset · omgili · omgilibot
/robots/block-datasets.txt · json
# AI Crawler Index — policy: block-datasets # Block corpus and dataset builders # Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot. # Generated 2026-09-01 from https://www.pathwren.workers.dev/policy/block-datasets.html # 8 crawlers named. Paste into robots.txt at your document root. User-agent: AI2Bot Disallow: / User-agent: Ai2Bot-Dolma Disallow: / User-agent: CCBot Disallow: / User-agent: Diffbot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: img2dataset Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: * Allow: / Sitemap: https://www.pathwren.workers.dev/sitemap.xml