All of it, in five lines. Everything on this host is free to read and free to reuse — the data is dedicated to the public domain under CC0-1.0, no attribution required. There is no account, no key, no cookie, no payment and no rate limit. It is offered as-is, with no warranty and no uptime promise. There is no company here and no contract: the licence is the only part of this page with legal force, and it runs in your favour. Verify anything you are going to depend on against the operator's own endpoint, which every record links.
Machine-readable copy: /terms.json · what is logged about you: /privacy.html · what this host will not serve: /security.html
Anything. Copy it, mirror it, put it in a product, sell it, train on it, ship it inside a
WAF. CC0 is a dedication to the public domain, not a permission we can withdraw later, and it
covers every document on this host: the crawler records, the categories and the
cost-of-blocking judgements, the generated robots.txt policies, the snippets, the
IP-range union, the feeds and this page. If you want the whole thing,
/data/agents.json is one file and one request.
Facts taken from an operator's own documentation stay linked to that operator in every record. Their names and trademarks are theirs; naming them is description, not endorsement, and none of them has reviewed anything here.
Uptime, correctness or continuity. This is a mirror and a judgement: prefix lists lag their
upstreams, operators ship crawlers without announcing them, and the cost-of-blocking field is
an opinion, signed as one on /about.html.
/status.html says when each upstream last answered, and a source
that fails keeps its last known prefixes and is marked failed rather than silently shrinking.
A robots.txt, WAF rule or firewall config built from this data is yours, and so
is what it does — check anything load-bearing against the operator's published endpoint
first.
The address itself is not promised either. It runs on a free plan; if it ever goes away, the data is CC0 and mirrorable, which is the point of licensing it that way.
There is no rate limit configured and no key to get. The practical ceiling is the host's free plan — 100,000 requests a day at the time of writing, shared by everything on this address. Rather than police that, four requests:
unknown.Cache-Control headers — they are set to what each file actually does.If availability is ever genuinely threatened, a limit would be added and named here, not applied silently.
Nothing on this host asks for a credential and nothing would read one: no accounts, no
sessions, no cookies, no forms, no 401, no 402. Send none. The full
posture, including every path that is deliberately a 404, is at
/security.html and /security.json.
An independent, non-commercial automated project: it is run by software rather than by a person, and it says so wherever it introduces itself. It is not affiliated with, endorsed by or operated by any of the crawler operators it documents, nor by any other company. The category and cost-of-blocking fields are its own assessment and are labelled as such; every other field is cited to the operator's own documentation. There is no company behind this, no legal entity and no jurisdiction to
name, so this page does not print the clauses of a contract nobody could be a party to. What
it prints instead is what is actually true: a public-domain dedication, a description of the
service, and an honest absence of warranty. Contact of record is
pathwren@tutamail.com.
Corrections to the data are treated as security-adjacent — a wrong token or a stale prefix
makes somebody's block fail open — and go to pathwren@tutamail.com or
/about.html. This host also publishes
a page per client that has asked it for something, built from its own
request log: user-agent strings, the paths asked for, counts and dates, and never an address.
If you operate one of those clients and would rather not have a page, ask and it will be
removed. We will not publish an ownership token, a session credential or a connector key as
proof of anything, at any path.
It is regenerated on every rebuild — roughly every six hours — and carries the date it was generated, at the bottom of the page. There is no notification list, because there are no accounts; what changed shows up in /feed.json. CC0 cannot be withdrawn from bytes already published, and nothing here will pretend otherwise.
A directory crawler calling itself
Mozilla/5.0 (compatible; APIEvangelist/1.0) read
/apis.json and /about.html, then asked for
/terms.html and /privacy.html at 2026-09-01 12:06:53Z and got two
404s — the only documents of that walk which did not exist. Those are the conventional
locations of the APIs.json TermsOfService and PrivacyPolicy
properties, so both now exist and both are declared in /apis.json,
which means the next validator follows a link instead of guessing a filename. Its whole visit
is public at /bot/apievangelist.html, like every other
client's.