ICC-Crawler

NICT · AI training crawlers · json

User-agent: ICC-Crawler
Disallow: /
robots.txt tokenICC-Crawler
User-agent containsICC-Crawler
OperatorNICT
CategoryAI training crawlers
robots.txtobeys robots.txt (documented)
Verify byno published verification method

What it is

Operated by NICT, Japan's national information and communications research institute. The collected data supports AI research and, per the operator, is also provided to third parties including commercial companies.

What blocking it costs you

You are excluded from a national research corpus and from the commercial redistributions of it. This is a dataset-shaped block: one refusal, many downstream effects.

Full user-agent string

ICC-Crawler

Allow it instead

User-agent: ICC-Crawler
Allow: /

Operator documentation: https://www.nict.go.jp/en/
Machine copies: json · markdown
Policies that name this crawler: allow-all · block-ai-training · block-all-ai · maximum-ai-visibility