CCBot

Common Crawl · Corpus and dataset builders · json

User-agent: CCBot
Disallow: /
robots.txt tokenCCBot
User-agent containsCCBot
OperatorCommon Crawl
CategoryCorpus and dataset builders
robots.txtobeys robots.txt (documented)
Verify byno published verification method

What it is

Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list.

What blocking it costs you

Future Common Crawl snapshots exclude you, so downstream training sets lose you too — but only going forward. Existing snapshots are permanent and blocking today does not retract them.

Full user-agent string

CCBot/2.0 (https://commoncrawl.org/faq/)

Allow it instead

User-agent: CCBot
Allow: /

Operator documentation: https://commoncrawl.org/faq
Machine copies: json · markdown
Policies that name this crawler: allow-all · block-all-ai · block-datasets · maximum-ai-visibility