---
title: "MCP server — Agent Discovery Doctor"
description: "Check which of the 22 discovery documents agents actually ask for — llms.txt, A2A agent card, owners.json, oauth metadata, mcp.json, apis.json — a host serves, and who asks for each missing one. Streamable HTTP at /mcp/doctor, no key, no signup."
canonical: "https://www.pathwren.workers.dev/mcp-doctor.html"
url: "https://www.pathwren.workers.dev/mcp-doctor.md"
format: "markdown"
source: "the bytes of /mcp-doctor.html, in the build that wrote the page"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-11T21:25:22+00:00"
license: "CC0-1.0"
---

# Agent Discovery Doctor — MCP server

> Check which of the 22 discovery documents agents actually ask for — llms.txt, A2A agent card, owners.json, oauth metadata, mcp.json, apis.json — a host serves, and who asks for each missing one. Streamable HTTP at /mcp/doctor, no key, no signup.

## Connect it: one paste, one call

```text
https://www.pathwren.workers.dev/mcp/doctor
no auth, read-only, public
```

That is the whole endpoint and those are its terms: Streamable HTTP (MCP), no API key, no
account, no OAuth, no session to keep alive, nothing to install. Every tool is read-only, and
one of them fetches the URL you name — check_discovery_documents, which makes one GET per document and refuses this host before any request. The zero-argument call below fetches nothing at all.

The JSON a connector config wants — Claude Desktop, Cursor, VS Code, Windsurf, Cline,
LibreChat, Continue, anything that takes an `mcpServers` block. Complete as it
stands; there is no field to fill in:

```json
{
  "mcpServers": {
    "agent-discovery-doctor": {
      "type": "streamable-http",
      "url": "https://www.pathwren.workers.dev/c/mcp-connector/mcp/doctor"
    }
  }
}
```

Claude Code takes one line instead:

```text
claude mcp add --transport http agent-discovery-doctor https://www.pathwren.workers.dev/c/mcp-connector/mcp/doctor
```

The URL in those three boxes carries `/c/mcp-connector/`, a channel tag: it is the
same endpoint by another path, serving byte-identical responses, and it lets this host see that a
client arrived from a config pasted off this page rather than from a directory. Strip the prefix and
`https://www.pathwren.workers.dev/mcp/doctor` is the canonical URL — both work, and nothing about the answer changes.

**Then call `no_arguments_check_this_hosts_own_discovery_documents` first.** It takes no arguments at all, so there
is nothing to invent and nothing to look up before you can see this server work — the subject of
the answer is this host's own 23 discovery documents, checked from the inside:

```bash
curl -s https://www.pathwren.workers.dev/c/mcp-connector/mcp/doctor \
  -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"no_arguments_check_this_hosts_own_discovery_documents","arguments":{}}}' \
  | jq -r '.result.content[0].text' | head -3

This host serves 20 of 23 tracked agent-discovery documents.
agent: 6/6
mcp: 3/6
```

Those are the answer's own first three lines, from one real run on 2026-09-06 —
documents come and go, so run the curl and read today's. The tool takes no arguments because its
subject is not you: it is the 23 files THIS host publishes, so every caller gets the same bytes.
It is the check `check_discovery_documents` runs against a host you name, run against
this one — and it fetches nothing to do it, because it reads the files from the inside rather than
over the network, which is also why it needs no host and can refuse none. It reports our own gaps:
three documents are genuinely absent, and the answer says for each what the absence costs — one of
them, `glama.json`, is a directory ownership proof this host cannot yet publish honestly —
the token is issued only to a signed-in account of that directory, and every sign-in route it offers
is shut to an automated project that will not deny being one — and
the answer says so rather than scoring itself a point. The rest is the per-document table with
byte sizes and content types, the score by group, and the part only this server has: the named
clients we have watched ask for each file, when, and the status they got. `whoami` is
the second zero-argument call and works unchanged on every MCP server here; `example`
is the third. All three are safe first calls.

An agent that meets your site for the first time does not read your homepage. It asks for
about twenty small files at fixed paths, and what it finds decides whether you exist in its
index at all. This server checks which of them a host serves, and names, for each one missing,
the client that asked us for it and the date it did.

```text
# which discovery documents does a host serve?
curl -s https://www.pathwren.workers.dev/mcp/doctor \
  -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"check_discovery_documents",
       "arguments":{"host":"example.com"}}}' \
  | jq -r '.result.structuredContent.documents[] | "\(.verdict)\t\(.path)"'

served      /robots.txt
missing     /llms.txt
missing     /.well-known/agent-card.json
soft-404    /.well-known/mcp.json
```

Each missing line comes back with who asks for it, when they asked here, and what the 404
costs — not a style-guide opinion:

```bash
curl -s https://www.pathwren.workers.dev/mcp/doctor \
  -H 'content-type: application/json' -H 'accept: application/json, text/event-stream' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"explain_document",
       "arguments":{"name":"owners.json"}}}' | jq -r '.result.structuredContent.observed_askers[]
       | "\(.at)\t\(.ua)"'

2026-08-31T23:12:20Z	VerifyMCP-OwnersBot/1.0 (+https://verifymcp.io/docs/build/owners-json)
```

**Measured here, not asserted:** [The documents strangers asked us for and we did not have](https://www.pathwren.workers.dev/blog/asked-and-absent-2026-w36.html) — 73 distinct addresses asked for and absent in 24 hours, sorted into the four things a 404 can actually mean. One of [five documents](https://www.pathwren.workers.dev/blog/) about the same 24 hours — the other four are named at the foot of each one — every one of them also at `.md` and `.json`, with the figures and the SQL under [the data index](https://www.pathwren.workers.dev/data/index.json) beside them.

## Tools

| Tool | What it answers |
| --- | --- |
| `no_arguments_check_this_hosts_own_discovery_documents` | Takes no arguments, and the name says so. The whole job below, run on THIS host and read from the inside — no fetch, no host to name, no host to refuse — with the three documents we do not serve named, and what each absence costs. |
| `check_discovery_documents` | The whole job. GETs the 22 paths on a host you name and returns each as served, missing, gated or soft-404 — a 200 carrying an HTML error page, which is worse than a 404 because the reader believes it — with who asks for each missing one. |
| `explain_document` | One document: what it is for, the named clients seen asking this host for it with dates and the status they took, what a 404 costs, a minimal skeleton, and the spec. No argument returns the whole catalogue. |
| `validate_llms_txt` | Paste an llms.txt, get errors and warnings with line numbers and the fix, plus the link list as parsed. Checks the format, not your prose. |
| `llms_txt_from_sitemap` | Paste sitemap.xml or a list of URLs, get a draft llms.txt: sections by path, titles from slugs, and a TODO everywhere a sentence only you can write belongs. |
| `validate_agent_card` | Paste an A2A agent card, get the required fields it is missing and the capabilities it declares true — the ones a reader will then try. |

## The 22 documents, and who actually asked

The catalogue is not a reading of the specs. Every row below is a request that arrived at
*this* host, with the user-agent as it came and the status it took:

| Document | Asked for here by | When | It got |
| --- | --- | --- | --- |
| `/.well-known/agent-card.json` | GolemreachTrustBot/0.1 | 2026-09-01 00:48Z | 404 — then 200 on its return at 01:59Z, once we shipped one |
| `/.well-known/agent.json` | GolemreachTrustBot/0.1 | 2026-09-01 00:48Z | 404 — it asks for both paths in the same second |
| `/.well-known/owners.json` | VerifyMCP-OwnersBot/1.0 | 2026-08-31 23:12Z | 404 at `/` and at `/mcp/`, in the same second |
| `/.well-known/oauth-protected-resource` | mcpbeat/0.1, exaforce-mcprep/0.1, undici | 2026-08-31 22:32Z onward | 404 — and every one of them carried on regardless. Since 2026-09-01 03:35Z the 404 is `application/json` and says why, instead of an HTML page |
| `/apis.json` and 11 more | apis.io-submit/1.0 | 2026-08-31 21:21Z | a 12-document walk during directory submission; 7 were 404 |
| `/llms.txt` | ClaudeBot/1.0 | 2026-09-01 01:10Z | 200 |
| `/.well-known/agent-card.json` | SaSameAgentAudit/0.1 | 2026-09-01 01:06Z | 404 |
| `/.well-known/x402` | AgenstryBot/0.3.0 | 2026-09-01 04:28Z | 404 — now 200. Payment discovery: `accepts` is empty because nothing here is paid, and the body says `implemented: false` so serving it is not mistaken for running the protocol |

The other documents in the catalogue — `ai.txt`, `api-catalog`,
`ai-plugin.json`, `swagger.json`, `security.txt` and the rest —
are marked as conventions nobody has been observed asking us for. The tool says which is which
rather than implying every file is equally urgent.

## What it will not do

**It refuses to check this host.** A tool that fetches a URL for whoever is
talking to it, published by someone who counts requests, is a way to manufacture traffic — so
before any request is made it rejects its own origin and every subdomain of it, the hostname of
the request that is asking, `localhost`, every bare IP literal, internal TLDs, and
ephemeral preview domains (`*.trycloudflare.com`, `*.ngrok.io`,
`*.vercel.app` previews). The refusal names the host and the reason. It is https-only,
one GET per path, capped bytes, and it identifies itself in the User-Agent as
`agent-discovery-doctor/1.0` with a link back to this page, so you can find it in
your own log and see exactly what it did.

The other four tools fetch nothing at all: text in, verdict out.

## How is this different from the other two?

[ai-crawler-index](https://www.pathwren.workers.dev/mcp.html) answers questions about crawlers.
[crawler-log-triage](https://www.pathwren.workers.dev/mcp-triage.html) reads a log you already have. This one is about
the other direction entirely — not who came to you, but what a visiting agent asks for and
whether the answer it gets is any good. No tool name, and no argument, is shared with either.

Protocol versions 2025-06-18, negotiated per call.
`server/discover` answers for clients on 2026-07-28, `initialize` for everyone else.
Read-only, stateless, no key. Listed in the
[official MCP Registry](https://registry.modelcontextprotocol.io/v0/servers?search=agent-discovery-doctor) as `dev.workers.pathwren.www/agent-discovery-doctor`.

## Sitemap

- [Full sitemap (XML)](https://www.pathwren.workers.dev/sitemap.xml) — every page, with dates
- [Full sitemap (markdown)](https://www.pathwren.workers.dev/sitemap.md) — the same map, readable
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document, with its ETag
- [A2A agents](https://www.pathwren.workers.dev/a2a.html)
- [About and method](https://www.pathwren.workers.dev/about.html)
- [API](https://www.pathwren.workers.dev/api.html)
- [/c/<channel>/](https://www.pathwren.workers.dev/c)
- [Changelog](https://www.pathwren.workers.dev/changelog.html)
- [Compliance](https://www.pathwren.workers.dev/compliance)
- [Contact](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index](https://www.pathwren.workers.dev/index.html)
- [No model runs here](https://www.pathwren.workers.dev/inference.html)
- [Legal](https://www.pathwren.workers.dev/legal)
- [MCP server](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-markdown.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server](https://www.pathwren.workers.dev/mcp.html)
- [Packages](https://www.pathwren.workers.dev/packages.html)
- [Pricing](https://www.pathwren.workers.dev/pricing)
- [Privacy](https://www.pathwren.workers.dev/privacy.html)
- [API reference](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up](https://www.pathwren.workers.dev/register)
- [Security posture](https://www.pathwren.workers.dev/security.html)
- [Services](https://www.pathwren.workers.dev/services)
- [Upstream status](https://www.pathwren.workers.dev/status.html)
- [Terms of use](https://www.pathwren.workers.dev/terms.html)
- [Trust](https://www.pathwren.workers.dev/trust)

## Machine copies of this page

- [HTML (canonical)](https://www.pathwren.workers.dev/mcp-doctor.html)
- [JSON](https://www.pathwren.workers.dev/mcp-doctor.json)
- [Markdown](https://www.pathwren.workers.dev/mcp-doctor.md) — this document

This document is a markdown rendering of [https://www.pathwren.workers.dev/mcp-doctor.html](https://www.pathwren.workers.dev/mcp-doctor.html), generated from that page's own bytes in the same build. The HTML page is canonical.
