# cohere-training-data-crawler

> Cohere's separately-named bulk crawler for model training data, split out so consent for training and consent for retrieval can differ.

| field | value |
|---|---|
| operator | Cohere |
| category | AI training crawlers |
| robots.txt token | `cohere-training-data-crawler` |
| user-agent contains | `cohere-training-data-crawler` |
| robots.txt | obeys robots.txt (documented) |
| verify by | no published verification method |
| published IP ranges | none published |
| prefixes mirrored | 0 IPv4 / 0 IPv6 |
| operator docs | https://cohere.com/ |
| last reviewed | 2026-09-01 |

## What blocking it costs you

Excluded from Cohere model training.

## Full user-agent string

```
Mozilla/5.0 (compatible; cohere-training-data-crawler)
```

## Block it

```
User-agent: cohere-training-data-crawler
Disallow: /
```

## Allow it

```
User-agent: cohere-training-data-crawler
Allow: /
```

JSON: https://www.pathwren.workers.dev/crawler/cohere-training-data-crawler.json · index: https://www.pathwren.workers.dev/llms.txt
