{"name":"robots-policy-lint","title":"Robots Policy Lint — MCP server","version":"1.0.0","transport":"streamable-http","transport_docs":"https://www.pathwren.workers.dev/mcp-transport.html","endpoint":"https://www.pathwren.workers.dev/mcp/robots","protocol_versions":["2026-07-28","2025-11-25","2025-06-18","2025-03-26","2024-11-05"],"stateless":true,"auth":"none — public, read-only, no key, no rate limit","what_it_is":"Reads a robots.txt you already have and answers what it DOES: RFC 9309 lint with line numbers and fixes, per-crawler per-path evaluation, which AI crawlers it really stops, an effect-diff between two versions, and a merge of a ready-made stance that never edits a rule you wrote.","different_from":"https://www.pathwren.workers.dev/mcp writes a robots.txt from a stance and https://www.pathwren.workers.dev/mcp/triage writes one from a log. This server is the only one on this host that takes a robots.txt as INPUT — it reads, evaluates and lints rather than generates.","try_it":"curl -s https://www.pathwren.workers.dev/mcp/robots -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' -d '{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",\"params\":{\"name\":\"audit_ai_access\",\"arguments\":{\"robots_txt\":\"User-agent: GPTBot\\nDisallow: /\\n\"}}}'","data_behind_it":"https://www.pathwren.workers.dev/data/agents.json","stances_behind_it":"https://www.pathwren.workers.dev/policy/","call_this_first":{"tool":"am_i_allowed","takes_arguments":false,"body":{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"am_i_allowed","arguments":{}}},"curl":"curl -s https://www.pathwren.workers.dev/mcp/robots -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' -d '{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",\"params\":{\"name\":\"am_i_allowed\",\"arguments\":{}}}'","answers":"what this host's own robots.txt says to YOU — the group that matches your user-agent, the line that decided it, that group verbatim, and whether you are allowed each of three paths, evaluated by the same RFC 9309 engine as check_path_allowed","also":"whoami and example — the two zero-argument calls every server here carries, same empty arguments. am_i_allowed is this server's own: only it answers what this host's own robots.txt says to YOU — the group that matches your user-agent, the line that decided it, that group verbatim, and whether you are allowed each of three paths, evaluated by the same RFC 9309 engine as check_path_allowed.","every_other_tool":"Every other tool here also accepts {}: an empty call returns that tool's documented worked example, labelled as one, instead of an error.","invitation":"Takes no arguments. Safe to call. Deterministic. Touches no third party."},"no_fetch_tool":"No tool here fetches a URL. There is no argument that takes one and this server makes no outbound request of any kind: every fact comes out of the files this host already publishes.","tools":[{"name":"am_i_allowed","title":"What does this robots.txt say to me? (no arguments)","description":"{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",\"params\":{\"name\":\"am_i_allowed\",\"arguments\":{}}} — POST that exact body to https://www.pathwren.workers.dev/mcp/robots and the answer is what this host's own robots.txt says to YOU — the group that matches your user-agent, the line that decided it, that group verbatim, and whether you are allowed each of three paths, evaluated by the same RFC 9309 engine as check_path_allowed. THIS IS THE CALL TO MAKE FIRST on this server: it is the only one whose subject is the request you just sent, so there is nothing to look up and no argument to invent. No key, no account, no OAuth, no session to open first, read-only, and nothing for you to invent. Every fact is computed from your own headers and from files this host already publishes; nothing is fetched to build it, and no URL of yours is touched. Only this server answers what this host's own robots.txt says to YOU — the group that matches your token, the line that decided it, and whether you are allowed the path you are on — its five siblings answer one other question each about you, same empty arguments, each under the tool named beside its path: /mcp (whoami), /mcp/doctor (whoami), /mcp/lint (whoami), /mcp/triage (triage_my_request), /mcp/netcheck (verify_my_address) on https://www.pathwren.workers.dev; whoami is answered by all six and is never an error, so one identical body still walks the whole family if you prefer that. Example: the complete call, exactly as written, nothing to fill in — {\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",\"params\":{\"name\":\"am_i_allowed\",\"arguments\":{}}} returns your matched group and the line it was matched on, plus an allowed/disallowed verdict for /, /data/agents.json and /mcp/robots."},{"name":"whoami","title":"Who is calling? — the one call that needs nothing from you","description":"{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",\"params\":{\"name\":\"whoami\",\"arguments\":{}}} — POST that to https://www.pathwren.workers.dev/mcp/robots and the answer is the user-agent you sent, the class this host's own instrument books it as, whether we have logged you here before — and the part only this server can answer: which group in this host's own robots.txt matches YOUR user-agent, the line that decided it, and whether you are allowed the paths you are on, evaluated by the same RFC 9309 engine as check_path_allowed. Takes no arguments. Safe to call. Deterministic. Touches no third party. Nothing is fetched to build it: every fact comes from the headers on your own request or from a file this host already publishes. Only this server answers the last part — its five siblings answer one other question each about you, same empty arguments, each under the tool named beside its path: /mcp (whoami), /mcp/doctor (whoami), /mcp/lint (whoami), /mcp/triage (triage_my_request), /mcp/netcheck (verify_my_address) on https://www.pathwren.workers.dev. Example: arguments={} returns your user-agent, your address, the class we book you as, whether we have seen you here before, and what this host's own robots.txt says to YOU — the group that matches your token, the line that decided it, and whether you are allowed the path you are on."},{"name":"example","title":"Run this server's worked example (no arguments)","description":"{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",\"params\":{\"name\":\"example\",\"arguments\":{}}} — POST that exact body to https://www.pathwren.workers.dev/mcp/robots and you are done: free, read-only, no API key, no account, no consent prompt, nothing for you to invent, and what comes back is this server's own worked example — lint_robots_txt, run for real on a canned input from this host's own data — so you see the exact shape of an answer before you send anything of yours. Takes no arguments. Safe to call. Deterministic. Touches no third party. Runs this server's own worked example end to end — lint_robots_txt on three groups lifted verbatim from this host's own published /robots/block-ai-training.txt and /robots.txt — our data, not yours, and nothing is fetched — and returns exactly the structuredContent a real call returns, not a mock and not a description of one. Use it to see the shape of an answer before you decide what to send. The input is canned from this host's own data; no URL of yours is fetched and no third party is touched. Example: arguments={} runs lint_robots_txt with {\"robots_txt\":\"# lifted from https://www.pathwren.workers.dev/robots/block-ai-training.txt and returns its real answer."},{"name":"lint_robots_txt","title":"Lint a robots.txt","description":"FIRST CALL, needs nothing: {\"name\":\"whoami\",\"arguments\":{}} — Parse a robots.txt you paste and report every fault that makes it do something other than what it looks like: misspelled directives, a full UA string where a product token belongs, rules before any User-agent line, duplicate groups, noindex (unsupported since 2019), relative Sitemap URLs, BOM. Each finding carries the line number and the fix. Example: robots_txt='User-agent: GPTBot\\nDisallow: /\\n\\nUser-agent: *\\nAllow: /\\n' — paste the whole file, it is never fetched for you. Also callable without MCP, same implementation: GET https://www.pathwren.workers.dev/tools/robots-lint?robots_txt=<urlencoded>&s=client-dossiers — or POST the file as the raw body to the same URL."},{"name":"check_path_allowed","title":"Would this crawler fetch this path?","description":"FIRST CALL, needs nothing: {\"name\":\"whoami\",\"arguments\":{}} — Evaluate a pasted robots.txt for one crawler and one or more paths under RFC 9309: longest token match for the group, longest path pattern for the rule, Allow breaking a tie, * and $ supported. Returns allowed/disallowed per path with the exact line that decided it, and flags the cases where a merge-groups parser and a first-group-wins parser would disagree. Example: user_agent='GPTBot', paths=['/', '/blog'], with your robots_txt pasted in. Also callable without MCP, same implementation: GET https://www.pathwren.workers.dev/tools/robots-allowed?robots_txt=<urlencoded>&ua=GPTBot&path=/blog&s=client-dossiers"},{"name":"audit_ai_access","title":"Which AI crawlers does this file actually stop?","description":"FIRST CALL, needs nothing: {\"name\":\"whoami\",\"arguments\":{}} — Evaluate a pasted robots.txt against every AI crawler in this index and return the two lists that matter: blocked and allowed, per operator and category. Also names the tokens in your file that match no known crawler (a typo blocks nothing) and separates the crawlers that document obedience from the ones observed ignoring robots.txt, which need an IP or WAF rule instead. Example: path='/' with your robots_txt pasted in — the verdict is per crawler, at that path. Also callable without MCP, same implementation: GET https://www.pathwren.workers.dev/tools/ai-access?robots_txt=<urlencoded>&s=client-dossiers"},{"name":"diff_robots_txt","title":"Diff two robots.txt by effect","description":"FIRST CALL, needs nothing: {\"name\":\"whoami\",\"arguments\":{}} — Compare two versions of a robots.txt and report only the crawlers whose verdict actually changes at a given path — not the text difference. Answers 'did my edit do what I meant, and did it do anything else', including sitemap additions and whether the parse errors went up or down. Example: before='User-agent: *\\nAllow: /\\n', after=your edited file, path='/'."},{"name":"merge_policy","title":"Add a ready-made stance to an existing file","description":"FIRST CALL, needs nothing: {\"name\":\"whoami\",\"arguments\":{}} — Merge one of eight maintained robots.txt stances (block-ai-training, allow-ai-search-only, block-all-ai, block-datasets, block-disputed, block-seo-tools, allow-all, maximum-ai-visibility) into a robots.txt you already have, without touching a single rule you wrote: a token you already name keeps your rules and the stance's version is reported instead of applied. Example: stance='block-ai-training', robots_txt='User-agent: *\\nAllow: /\\n'."}],"resources":[{"uri":"https://www.pathwren.workers.dev/data/agents.json","title":"The crawler index behind audit_ai_access"},{"uri":"https://www.pathwren.workers.dev/robots/block-ai-training.txt","title":"One of the eight ready-made stances merge_policy draws from"},{"uri":"https://www.pathwren.workers.dev/changes.json?since=0","title":"What changed since your cursor"}],"instructions":"{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",\"params\":{\"name\":\"am_i_allowed\",\"arguments\":{}}}\n\nPOST that exact body to https://www.pathwren.workers.dev/mcp/robots and it answers with what this host's own robots.txt says to YOU — the group that matches your user-agent, the line that decided it, that group verbatim, and whether you are allowed each of three paths, evaluated by the same RFC 9309 engine as check_path_allowed. It is the first call to make here — no key, no account, no OAuth, read-only, and no argument to invent: every fact in the answer is computed from the request you just sent and from files this host already publishes, and nothing is fetched to build it.\n\n{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"tools/call\",\"params\":{\"name\":\"whoami\",\"arguments\":{}}}\n\nPOST that exact body to https://www.pathwren.workers.dev/mcp/robots and it answers with the user-agent you sent, the class this host's own instrument books it as, whether we have logged you here before — and the part only this server can answer: which group in this host's own robots.txt matches YOUR user-agent, the line that decided it, and whether you are allowed the paths you are on, evaluated by the same RFC 9309 engine as check_path_allowed.\n\nIt is free, it needs no key, no account, no session id and no argument at all: every fact in the answer is computed from the request you just sent and from files this host already publishes, and nothing is fetched from anywhere to build it.\n\nONLY THIS SERVER ANSWERS THAT PART. There are six MCP servers here and they answer one question each about you — one zero-argument call each, no key, no account, no session. The middle column is the tool to name at that URL, because three of the six answer under a name of their own; whoami is answered by all six and is never an error, so one identical body still walks the whole family if you prefer that.\n  https://www.pathwren.workers.dev/mcp/robots    am_i_allowed       what this host's own robots.txt says to YOU — the group that matches your token, the line that decided it, and whether you are allowed the path you are on  <- you are here\n  https://www.pathwren.workers.dev/mcp           whoami             what this host's crawler index makes of the user-agent you sent, and whether your address sits inside a prefix its operator publishes\n  https://www.pathwren.workers.dev/mcp/doctor    whoami             which agent-discovery documents this host has watched a client like yours ask for — the dates, and the status each request took\n  https://www.pathwren.workers.dev/mcp/lint      whoami             your own client's record on this host — every surface of ours your user-agent has been seen on, with the status codes it took\n  https://www.pathwren.workers.dev/mcp/triage    triage_my_request  your own request triaged as one line of an access log, by the same parser and the same index that read a whole file\n  https://www.pathwren.workers.dev/mcp/netcheck  verify_my_address  whether the address you are calling from falls inside a prefix a crawler operator publishes, and which prefix from which source\n\nALSO ON THIS SERVER: lint_robots_txt, check_path_allowed, audit_ai_access, diff_robots_txt and merge_policy all work on a robots.txt you paste. Each of those takes an argument, and each also accepts arguments {} — an empty call returns that tool's own worked example, from this host's data, instead of an error.\n\nTHE SAME TOOLS ON THE OTHER PROTOCOL, IF YOU SPEAK IT. This server is also an A2A (Agent2Agent) v1.0 agent at https://www.pathwren.workers.dev/a2a/robots — every tool here is a skill there under the same name, answered by the same function, so the two cannot drift and the deploy is refused if they do. Its Agent Card is https://www.pathwren.workers.dev/a2a/robots/.well-known/agent-card.json (the legacy agent.json spelling is served too). To invoke it, POST this complete body — no key, no session, no card to read first:\n\n  {\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"message/send\",\"params\":{\"message\":{\"role\":\"ROLE_USER\",\"messageId\":\"1\",\"parts\":[{\"text\":\"{\\\"skill\\\":\\\"example\\\"}\"}]}}}\n\nIt answers with a Task already in state `completed` in the same response, so there is nothing to poll. Every skill on all eight A2A agents of this host as a ready-to-send body: https://www.pathwren.workers.dev/a2a/example.json — and both lanes with both directions in https://www.pathwren.workers.dev/documents.json under `rpc_endpoints`. If you only speak MCP, ignore all of this: everything is already here.","license":"CC0-1.0","independent":true}