curl -s https://www.pathwren.workers.dev/crawler/index.json   # this page, as JSON
curl -s https://www.pathwren.workers.dev/data/agents.json     # Every crawler record in one file

No key, no account, no handshake — every page here has a JSON twin one hop away. Machine doors: 6 keyless GET tools · documents.json · changes · llms.txt · openapi.json · agent card · mcp · a2a

All crawlers (150)

Grouped by what the crawl is for. Machine copy: agents.json · agents.csv · user-agents.txt

Measured, not asserted: 24 hours of crawler and agent traffic against this host — 2,921 unique clients, 1,229 of them crawlers, and 72.4% of the busiest hour was one bot from one address. Dated 2026-09-06, one named window, every figure with its SQL at /data/traffic-2026-w36.json.

Measured here, not asserted: 24 hours of AI-crawler traffic, as the actual table — 3,374 client keys, 1,313 distinct user-agent strings and the AS organisation behind every one of them, in the 24 hours to 2026-09-06T16:00:40+00:00. Published as the per-UA and per-ASN table rather than as a claim about it, with the SQL beside it. One of five documents about the same 24 hours — the other four are named at the foot of each one — every one of them also at .md and .json, with the figures and the SQL under the data index beside them.

AI training crawlers (27)

Collect pages in bulk so that a model can be trained or fine-tuned on them. Blocking these removes you from future training sets and changes nothing a user sees today.

Crawlerrobots.txt tokenOperatorrobots.txt
anthropic-aianthropic-aiAnthropicn-a
Applebot-ExtendedApplebot-ExtendedApplen-a
BytespiderBytespiderByteDancedisputed
ClaudeBotClaudeBotAnthropicdocumented
cohere-training-data-crawlercohere-training-data-crawlerCoheredocumented
CotoyogiCotoyogiROIS-DSdocumented
FacebookBotFacebookBotMetadocumented
Factset_spyderbotFactset_spyderbotFactSetundocumented
Google-ExtendedGoogle-ExtendedGooglen-a
GoogleOtherGoogleOtherGoogledocumented
GoogleOther-ImageGoogleOther-ImageGoogledocumented
GoogleOther-VideoGoogleOther-VideoGoogledocumented
GPTBotGPTBotOpenAIdocumented
ICC-CrawlerICC-CrawlerNICTdocumented
ISSCyberRiskCrawlerISSCyberRiskCrawlerISS Corporate Solutionsdisputed
Linguee BotLinguee BotLingueedisputed
meta-externalagentmeta-externalagentMetadocumented
Poseidon Research CrawlerPoseidon Research CrawlerPoseidon Researchundocumented
QuillBotQuillBotQuillBotundocumented
ReflectionbotReflectionbotReflection AIundocumented
SBIntuitionsBotSBIntuitionsBotSB Intuitionsdocumented
SemrushBot-OCOBSemrushBot-OCOBSemrushdocumented
Sidetrade indexer botSidetrade indexer botSidetradeundocumented
TikTokSpiderTikTokSpiderByteDancedisputed
Webzio-ExtendedWebzio-ExtendedWebz.iodocumented
YandexAdditionalYandexAdditionalYandexown-token-only
YandexAdditionalBotYandexAdditionalBotYandexown-token-only

Build the retrieval index an assistant answers and cites from. These are the crawlers that send you traffic; blocking them is the expensive mistake in this space.

Crawlerrobots.txt tokenOperatorrobots.txt
AIWebIndexAIWebIndexLyrenthdocumented
AmazonbotAmazonbotAmazondocumented
AndibotAndibotAndiundocumented
AnomuraAnomuraDireqtdocumented
atlassian-botatlassian-botAtlassiandocumented
bedrockbotbedrockbotAmazondocumented
Claude-SearchBotClaude-SearchBotAnthropicdocumented
Claude-WebClaude-WebAnthropicn-a
Cloudflare-AutoRAGCloudflare-AutoRAGCloudflaredocumented
DuckAssistBotDuckAssistBotDuckDuckGodocumented
ExaSearchBotExaSearchBotExaundocumented
Google-CloudVertexBotGoogle-CloudVertexBotGoogledocumented
KlaviyoAIBotKlaviyoAIBotKlaviyodocumented
Meta-WebIndexerMeta-WebIndexerMetaundocumented
OAI-SearchBotOAI-SearchBotOpenAIdocumented
PerplexityBotPerplexityBotPerplexitydocumented
PhindBotPhindBotPhindundocumented
QualifiedBotQualifiedBotQualifiedundocumented
ShapBotShapBotParalleldocumented
TerraCottaTerraCottaCeramic AIdocumented
YouBotYouBotYou.comdocumented

User-triggered fetchers (12)

Fetch one page because a person asked for it, right then. One human intent, one request. Blocking them produces a visible error for a real reader.

Crawlerrobots.txt tokenOperatorrobots.txt
ChatGPT AgentChatGPT-UserOpenAIdocumented
ChatGPT-UserChatGPT-UserOpenAIdocumented
Claude-UserClaude-UserAnthropicdocumented
cohere-aicohere-aiCoheredocumented
Google-AgentGoogle-AgentGoogleby-design-no
Google-GeminiNotebookGoogle-GeminiNotebookGoogleby-design-no
Google-PinpointGoogle-PinpointGoogleby-design-no
Google-Read-AloudGoogle-Read-AloudGoogleby-design-no
meta-externalfetchermeta-externalfetcherMetadocumented
MistralAI-UserMistralAI-UserMistral AIdocumented
Perplexity-UserPerplexity-UserPerplexityby-design-no
YandexCalendarYandexCalendarYandexown-token-only

Corpus and dataset builders (17)

Crawl the web into a published or resold dataset that other people train on. Highest leverage per block, longest delay before any effect.

Crawlerrobots.txt tokenOperatorrobots.txt
AI2BotAI2BotAllen Institute for AIdocumented
Ai2Bot-DolmaAi2Bot-DolmaAllen Institute for AIdocumented
aiHitBotaiHitBotaiHitdocumented
AwarioRssBotAwarioRssBotAwariodocumented
AwarioSmartBotAwarioSmartBotAwariodocumented
CCBotCCBotCommon Crawldocumented
DiffbotDiffbotDiffbotdocumented
EchoboxBotEchoboxBotEchoboxundocumented
ImagesiftBotImagesiftBotHive AIdocumented
img2datasetimg2datasetLAION / img2datasetdocumented
LAIONDownloaderLAIONDownloaderLAION / img2datasetby-design-no
omgiliomgiliWebz.iodocumented
omgilibotomgilibotWebz.iodocumented
Panscientpanscient.comPanscientdocumented
ThinkbotThinkbotThinkbotdisputed
VelenPublicWebCrawlerVelenPublicWebCrawlerHunter (Velen)documented
YaKYaKMeltwaterundocumented

Classic index-and-rank crawlers. Several also feed their operator's generative answers, which is why the AI opt-out for Google and Apple is a token rather than a block.

Crawlerrobots.txt tokenOperatorrobots.txt
ApplebotApplebotAppledocumented
BaiduspiderBaiduspiderBaidudocumented
bingbotbingbotMicrosoftdocumented
DuckDuckBotDuckDuckBotDuckDuckGodocumented
GooglebotGooglebotGoogledocumented
Googlebot-ImageGooglebot-ImageGoogledocumented
Googlebot-NewsGooglebot-NewsGoogledocumented
Googlebot-VideoGooglebot-VideoGoogledocumented
KagibotKagibotKagidocumented
MojeekBotMojeekBotMojeekdocumented
PetalBotPetalBotHuaweidocumented
PinterestbotPinterestbotPinterestdocumented
QwantbotQwantbotQwantdocumented
Qwantbot-newsQwantbot-newsQwantdocumented
SeznamBotSeznamBotSeznamdocumented
Storebot-GoogleStorebot-GoogleGoogledocumented
TimpibotTimpibotTimpidocumented
YandexBlogsYandexBlogsYandexdocumented
YandexBotYandexBotYandexdocumented
YandexComBotYandexComBotYandexown-token-only
YandexFaviconsYandexFaviconsYandexown-token-only
YandexImagesYandexImagesYandexdocumented
YandexMarketYandexMarketYandexdocumented
YandexMediaYandexMediaYandexdocumented
YandexMobileBotYandexMobileBotYandexown-token-only
YandexRenderResourcesBotYandexRenderResourcesBotYandexown-token-only
YandexVideoYandexVideoYandexdocumented
YetiYetiNaverdocumented

SEO and backlink crawlers (17)

Commercial link-graph tooling. No user-facing effect either way, and usually a large share of your bot bandwidth.

Crawlerrobots.txt tokenOperatorrobots.txt
AhrefsBotAhrefsBotAhrefsdocumented
AhrefsSiteAuditAhrefsSiteAuditAhrefsdocumented
BarkrowlerbarkrowlerBabbardocumented
DataForSeoBotDataForSeoBotDataForSEOdocumented
DotBotdotbotMozdocumented
MJ12botMJ12botMajesticdocumented
rogerbotrogerbotMozdocumented
SemrushBotSemrushBotSemrushdocumented
SemrushBot-BASemrushBot-BASemrushdocumented
SemrushBot-ESISemrushBot-ESISemrushdocumented
SemrushBot-FTSemrushBot-FTSemrushdocumented
SemrushBot-SISemrushBot-SISemrushdocumented
SemrushBot-SWASemrushBot-SWASemrushdocumented
SEOkicksSEOkicksSEOkicksdocumented
serpstatbotserpstatbotSerpstatdocumented
SiteAuditBotSiteAuditBotSemrushdocumented
SplitSignalBotSplitSignalBotSemrushdocumented

Archivers (2)

Preservation crawlers. Their output is public and permanent, which makes them a separate decision from the AI one.

Crawlerrobots.txt tokenOperatorrobots.txt
archive.org_botarchive.org_botInternet Archivedocumented
ia_archiveria_archiverInternet Archivedocumented

Link preview fetchers (4)

Read your Open Graph tags when someone shares a link. Blocking these is almost always an accident.

Crawlerrobots.txt tokenOperatorrobots.txt
facebookexternalhitfacebookexternalhitMetadocumented
GoogleMessagesGoogleMessagesGoogleby-design-no
SlackbotSlackbotSlackdocumented
Slackbot-LinkExpandingSlackbot-LinkExpandingSlackdocumented

Tools and frameworks (22)

Not operators: crawling software anyone can run. The party behind the request is unknown, so treat them as a rate-limit question rather than a consent question.

Crawlerrobots.txt tokenOperatorrobots.txt
AdsBot-GoogleAdsBot-GoogleGoogleown-token-only
AdsBot-Google-MobileAdsBot-Google-MobileGoogleown-token-only
AdsBot-Google-Mobile-AppsAdsBot-Google-Mobile-AppsGoogleown-token-only
APIs-GoogleAPIs-GoogleGoogleown-token-only
Crawl4AICrawl4AICrawl4AI projectundocumented
CrawlspaceCrawlspaceCrawlspacedocumented
FeedFetcher-GoogleFeedFetcher-GoogleGoogleby-design-no
FirecrawlAgentFirecrawlAgentFirecrawldocumented
Google-CWSGoogle-CWSGoogleby-design-no
Google-InspectionToolGoogle-InspectionToolGoogledocumented
Google-SafetyGoogle-SafetyGoogleby-design-no
Google-Site-VerificationGoogle-Site-VerificationGoogleby-design-no
GoogleProducerGoogleProducerGoogleby-design-no
LightpandaLightpandaLightpandaundocumented
Mediapartners-GoogleMediapartners-GoogleGoogleown-token-only
ScrapyScrapyScrapy projectdocumented
Screaming Frog SEO SpiderScreaming Frog SEO SpiderScreaming Frogdocumented
wpbotwpbotQuantumCloudundocumented
YandexDirectYandexDirectYandexown-token-only
YandexMetrikaYandexMetrikaYandexby-design-no
YandexScreenshotBotYandexScreenshotBotYandexown-token-only
YandexWebmasterYandexWebmasterYandexdocumented

This page as markdown: /c/mbin/crawler/index.md — the same text, no markup to strip, no JavaScript, no key, CC0. Every page here has one: add .md to any address (also .mdx, <page>.html.md, <page>.html.mdx), or send Accept: text/markdown to this one. All of them in a single index: /c/mbin/sitemap.md.