Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect.
| Crawler | Token | Operator | Cost of blocking |
|---|---|---|---|
| AI2Bot | AI2Bot | Allen Institute for AI | Excluded from open research datasets. Worth a deliberate decision: this is the category wh… |
| Ai2Bot-Dolma | Ai2Bot-Dolma | Allen Institute for AI | Same as AI2Bot: exclusion from an open, published training corpus.… |
| CCBot | CCBot | Common Crawl | Future Common Crawl snapshots exclude you, so downstream training sets lose you too — but … |
| Diffbot | Diffbot | Diffbot | Your facts stop entering a widely-licensed knowledge graph. Whether that is a loss depends… |
| ImagesiftBot | ImagesiftBot | Hive AI | Your images stop entering an image dataset and reverse-image index.… |
| img2dataset | img2dataset | LAION / img2dataset | Your images are skipped when someone materialises an image-text dataset that references th… |
| omgili | omgili | Webz.io | Same as omgilibot.… |
| omgilibot | omgilibot | Webz.io | Exclusion from a commercial dataset resold to third parties.… |