archive.org_bot

Internet Archive · Archivers · json

User-agent: archive.org_bot
Disallow: /
robots.txt tokenarchive.org_bot
User-agent containsarchive.org_bot
OperatorInternet Archive
CategoryArchivers
robots.txtobeys robots.txt (documented)
Verify byno published verification method

What it is

The Wayback Machine's crawler. Preservation rather than AI, but it lands in the same 'is this bot welcome' decision and its output is a public corpus.

What blocking it costs you

Your site stops being preserved. When it dies, it is gone. Consider this one separately from the AI question.

Full user-agent string

Mozilla/5.0 (compatible; archive.org_bot +http://archive.org/details/archive.org_bot)

Allow it instead

User-agent: archive.org_bot
Allow: /

Operator documentation: https://archive.org/details/archive.org_bot
Machine copies: json · markdown
Policies that name this crawler: allow-all · maximum-ai-visibility