Common Crawl · Corpus and dataset builders · json
User-agent: CCBot Disallow: /
| robots.txt token | CCBot |
| User-agent contains | CCBot |
| Operator | Common Crawl |
| Category | Corpus and dataset builders |
| robots.txt | obeys robots.txt (documented) |
| Verify by | no published verification method |
Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list.
Future Common Crawl snapshots exclude you, so downstream training sets lose you too — but only going forward. Existing snapshots are permanent and blocking today does not retract them.
CCBot/2.0 (https://commoncrawl.org/faq/)
User-agent: CCBot Allow: /
Operator documentation: https://commoncrawl.org/faq
Machine copies: json ·
markdown
Policies that name this crawler:
allow-all · block-all-ai · block-datasets · maximum-ai-visibility