Block corpus and dataset builders

Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot.

curl -s https://www.pathwren.workers.dev/robots/block-datasets.txt >> robots.txt

These are the highest-leverage blocks per line, because one crawl becomes many downstream training runs. It is also the block with the longest delay before it has any effect, and no effect at all on archives already published.

Names 8 crawlers

AI2Bot · Ai2Bot-Dolma · CCBot · Diffbot · ImagesiftBot · img2dataset · omgili · omgilibot

The file

/robots/block-datasets.txt · json

# AI Crawler Index — policy: block-datasets
# Block corpus and dataset builders
# Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot.
# Generated 2026-09-01 from https://www.pathwren.workers.dev/policy/block-datasets.html
# 8 crawlers named. Paste into robots.txt at your document root.

User-agent: AI2Bot
Disallow: /

User-agent: Ai2Bot-Dolma
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: ImagesiftBot
Disallow: /

User-agent: img2dataset
Disallow: /

User-agent: omgili
Disallow: /

User-agent: omgilibot
Disallow: /

User-agent: *
Allow: /

Sitemap: https://www.pathwren.workers.dev/sitemap.xml