# ICC-Crawler

> Operated by NICT, Japan's national information and communications research institute. The collected data supports AI research and, per the operator, is also provided to third parties including commercial companies.

| field | value |
|---|---|
| operator | NICT |
| category | AI training crawlers |
| robots.txt token | `ICC-Crawler` |
| user-agent contains | `ICC-Crawler` |
| robots.txt | obeys robots.txt (documented) |
| verify by | no published verification method |
| published IP ranges | none published |
| prefixes mirrored | 0 IPv4 / 0 IPv6 |
| operator docs | https://www.nict.go.jp/en/ |
| last reviewed | 2026-09-01 |

## What blocking it costs you

You are excluded from a national research corpus and from the commercial redistributions of it. This is a dataset-shaped block: one refusal, many downstream effects.

## Full user-agent string

```
ICC-Crawler
```

## Block it

```
User-agent: ICC-Crawler
Disallow: /
```

## Allow it

```
User-agent: ICC-Crawler
Allow: /
```

JSON: https://www.pathwren.workers.dev/crawler/icc-crawler.json · index: https://www.pathwren.workers.dev/llms.txt
