---
title: "Sitemap — AI Crawler Index"
description: "Every page AI Crawler Index publishes, grouped by section, with the machine copy of each."
canonical: "https://www.pathwren.workers.dev/sitemap.md"
url: "https://www.pathwren.workers.dev/sitemap.md"
format: "markdown"
generator: "surfaces/ai-crawler-index/build.py"
generated: "2026-09-06T07:19:53+00:00"
license: "CC0-1.0"
---

# Sitemap

> Every page AI Crawler Index publishes. 302 pages, grouped by section. Every page also answers as JSON (`.json`) and as markdown (`.md`) at the same address — and at `.mdx`, `<page>.html.md` and `<page>.html.mdx`, or by sending `Accept: text/markdown` to the page itself. Same bytes, no markup to strip, no JavaScript, no key. The HTML page stays canonical and every mirror says so in a `Link:` header.

## Maps of this host

- [sitemap.xml](https://www.pathwren.workers.dev/sitemap.xml) — every page with its last-modified date
- [llms.txt](https://www.pathwren.workers.dev/llms.txt) — the whole host in one text file
- [llms-full.txt](https://www.pathwren.workers.dev/llms-full.txt) — the same, with the content inlined
- [documents.json](https://www.pathwren.workers.dev/documents.json) — every document with its strong ETag
- [openapi.json](https://www.pathwren.workers.dev/openapi.json) — every read endpoint, described formally

## Pages

- [A2A agents — AI Crawler Index](https://www.pathwren.workers.dev/a2a.html)
- [About and method — AI Crawler Index](https://www.pathwren.workers.dev/about.html)
- [API — AI Crawler Index](https://www.pathwren.workers.dev/api.html)
- [/c/<channel>/ — channel-tagged copies — AI Crawler Index](https://www.pathwren.workers.dev/c)
- [Changelog — AI Crawler Index](https://www.pathwren.workers.dev/changelog.html)
- [Compliance — AI Crawler Index](https://www.pathwren.workers.dev/compliance)
- [Contact — AI Crawler Index](https://www.pathwren.workers.dev/contact)
- [Impressum · Anbieterkennzeichnung — AI Crawler Index](https://www.pathwren.workers.dev/impressum)
- [AI Crawler Index — 150 AI crawlers, their robots.txt tokens and IP ranges](https://www.pathwren.workers.dev/index.html)
- [No model runs here — /chat/completions and the rest, refused](https://www.pathwren.workers.dev/inference.html)
- [Legal — AI Crawler Index](https://www.pathwren.workers.dev/legal)
- [MCP server — Agent Discovery Doctor](https://www.pathwren.workers.dev/mcp-doctor.html)
- [MCP server — MCP Endpoint Lint](https://www.pathwren.workers.dev/mcp-lint.html)
- [MCP server — Crawler IP Verifier](https://www.pathwren.workers.dev/mcp-netcheck.html)
- [MCP server — Robots Policy Lint](https://www.pathwren.workers.dev/mcp-robots.html)
- [MCP transport: the GET and HEAD leg](https://www.pathwren.workers.dev/mcp-transport.html)
- [MCP server — Crawler Log Triage](https://www.pathwren.workers.dev/mcp-triage.html)
- [MCP server — AI Crawler Index](https://www.pathwren.workers.dev/mcp.html)
- [Packages — AI Crawler Index](https://www.pathwren.workers.dev/packages.html)
- [Pricing — AI Crawler Index](https://www.pathwren.workers.dev/pricing)
- [Privacy — AI Crawler Index](https://www.pathwren.workers.dev/privacy.html)
- [API reference — AI Crawler Index](https://www.pathwren.workers.dev/reference)
- [Access, keys and sign-up — AI Crawler Index](https://www.pathwren.workers.dev/register)
- [Security posture — AI Crawler Index](https://www.pathwren.workers.dev/security.html)
- [Services — AI Crawler Index](https://www.pathwren.workers.dev/services)
- [Upstream status — AI Crawler Index](https://www.pathwren.workers.dev/status.html)
- [Terms of use — AI Crawler Index](https://www.pathwren.workers.dev/terms.html)
- [Trust — AI Crawler Index](https://www.pathwren.workers.dev/trust)

## ai-crawler-logs

- [ai-crawler-logs — Who was actually in your access log, and what to paste to act on it.](https://www.pathwren.workers.dev/ai-crawler-logs/index.html)

## ai-crawler-robots

- [ai-crawler-robots — Does your robots.txt block the crawlers you think it blocks?](https://www.pathwren.workers.dev/ai-crawler-robots/index.html)

## blog

- [Your robots.txt is probably blocking the wrong AI crawlers — AI Crawler Index](https://www.pathwren.workers.dev/blog/ai-crawler-cost.html)
- [AI crawler traffic, one host, 24 hours: 2,921 clients — and the busiest hour of the day was a single bot — AI Crawler Index](https://www.pathwren.workers.dev/blog/ai-crawler-traffic-2026-w36.html)
- [I logged every client that hit my site for 24 hours: 1,095 of them, and 179 claimed to be people — AI Crawler Index](https://www.pathwren.workers.dev/blog/client-census.html)
- [Blog — AI Crawler Index](https://www.pathwren.workers.dev/blog/index.html)

## category

- [AI search crawlers (21) — AI Crawler Index](https://www.pathwren.workers.dev/category/ai-search.html)
- [AI training crawlers (27) — AI Crawler Index](https://www.pathwren.workers.dev/category/ai-training.html)
- [Archivers (2) — AI Crawler Index](https://www.pathwren.workers.dev/category/archive.html)
- [Corpus and dataset builders (17) — AI Crawler Index](https://www.pathwren.workers.dev/category/dataset.html)
- [Link preview fetchers (4) — AI Crawler Index](https://www.pathwren.workers.dev/category/preview.html)
- [Search engines (28) — AI Crawler Index](https://www.pathwren.workers.dev/category/search.html)
- [SEO and backlink crawlers (17) — AI Crawler Index](https://www.pathwren.workers.dev/category/seo.html)
- [Tools and frameworks (22) — AI Crawler Index](https://www.pathwren.workers.dev/category/tool.html)
- [User-triggered fetchers (12) — AI Crawler Index](https://www.pathwren.workers.dev/category/user-fetch.html)

## crawler

- [AdsBot-Google-Mobile-Apps — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/adsbot-google-mobile-apps.html)
- [AdsBot-Google-Mobile — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/adsbot-google-mobile.html)
- [AdsBot-Google — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/adsbot-google.html)
- [AhrefsBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/ahrefsbot.html)
- [AhrefsSiteAudit — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/ahrefssiteaudit.html)
- [Ai2Bot-Dolma — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/ai2bot-dolma.html)
- [AI2Bot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/ai2bot.html)
- [aiHitBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/aihitbot.html)
- [AIWebIndex — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/aiwebindex.html)
- [Amazonbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/amazonbot.html)
- [Andibot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/andibot.html)
- [Anomura — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/anomura.html)
- [anthropic-ai — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/anthropic-ai.html)
- [APIs-Google — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/apis-google.html)
- [Applebot-Extended — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/applebot-extended.html)
- [Applebot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/applebot.html)
- [archive.org_bot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/archive-org-bot.html)
- [atlassian-bot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/atlassian-bot.html)
- [AwarioRssBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/awariorssbot.html)
- [AwarioSmartBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/awariosmartbot.html)
- [Baiduspider — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/baiduspider.html)
- [Barkrowler — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/barkrowler.html)
- [bedrockbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/bedrockbot.html)
- [bingbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/bingbot.html)
- [Bytespider — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/bytespider.html)
- [CCBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/ccbot.html)
- [ChatGPT Agent — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/chatgpt-agent.html)
- [ChatGPT-User — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/chatgpt-user.html)
- [Claude-SearchBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/claude-searchbot.html)
- [Claude-User — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/claude-user.html)
- [Claude-Web — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/claude-web.html)
- [ClaudeBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/claudebot.html)
- [Cloudflare-AutoRAG — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/cloudflare-autorag.html)
- [cohere-ai — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/cohere-ai.html)
- [cohere-training-data-crawler — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/cohere-training-data-crawler.html)
- [Cotoyogi — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/cotoyogi.html)
- [Crawl4AI — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/crawl4ai.html)
- [Crawlspace — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/crawlspace.html)
- [DataForSeoBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/dataforseobot.html)
- [Diffbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/diffbot.html)
- [DotBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/dotbot.html)
- [DuckAssistBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/duckassistbot.html)
- [DuckDuckBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/duckduckbot.html)
- [EchoboxBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/echoboxbot.html)
- [ExaSearchBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/exasearchbot.html)
- [FacebookBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/facebookbot.html)
- [facebookexternalhit — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/facebookexternalhit.html)
- [Factset_spyderbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/factset-spyderbot.html)
- [FeedFetcher-Google — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/feedfetcher-google.html)
- [FirecrawlAgent — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/firecrawlagent.html)
- [Google-Agent — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-agent.html)
- [Google-CloudVertexBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-cloudvertexbot.html)
- [Google-CWS — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-cws.html)
- [Google-Extended — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-extended.html)
- [Google-GeminiNotebook — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-gemininotebook.html)
- [Google-InspectionTool — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-inspectiontool.html)
- [Google-Pinpoint — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-pinpoint.html)
- [Google-Read-Aloud — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-read-aloud.html)
- [Google-Safety — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-safety.html)
- [Google-Site-Verification — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/google-site-verification.html)
- [Googlebot-Image — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googlebot-image.html)
- [Googlebot-News — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googlebot-news.html)
- [Googlebot-Video — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googlebot-video.html)
- [Googlebot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googlebot.html)
- [GoogleMessages — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googlemessages.html)
- [GoogleOther-Image — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googleother-image.html)
- [GoogleOther-Video — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googleother-video.html)
- [GoogleOther — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googleother.html)
- [GoogleProducer — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/googleproducer.html)
- [GPTBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/gptbot.html)
- [ia_archiver — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/ia-archiver.html)
- [ICC-Crawler — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/icc-crawler.html)
- [ImagesiftBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/imagesiftbot.html)
- [img2dataset — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/img2dataset.html)
- [All 150 crawlers — AI Crawler Index](https://www.pathwren.workers.dev/crawler/index.html)
- [ISSCyberRiskCrawler — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/isscyberriskcrawler.html)
- [Kagibot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/kagibot.html)
- [KlaviyoAIBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/klaviyoaibot.html)
- [LAIONDownloader — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/laiondownloader.html)
- [Lightpanda — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/lightpanda.html)
- [Linguee Bot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/linguee-bot.html)
- [Mediapartners-Google — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/mediapartners-google.html)
- [meta-externalagent — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/meta-externalagent.html)
- [meta-externalfetcher — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/meta-externalfetcher.html)
- [Meta-WebIndexer — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/meta-webindexer.html)
- [MistralAI-User — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/mistralai-user.html)
- [MJ12bot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/mj12bot.html)
- [MojeekBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/mojeekbot.html)
- [OAI-SearchBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/oai-searchbot.html)
- [omgili — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/omgili.html)
- [omgilibot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/omgilibot.html)
- [Panscient — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/panscient.html)
- [Perplexity-User — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/perplexity-user.html)
- [PerplexityBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/perplexitybot.html)
- [PetalBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/petalbot.html)
- [PhindBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/phindbot.html)
- [Pinterestbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/pinterestbot.html)
- [Poseidon Research Crawler — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/poseidon-research-crawler.html)
- [QualifiedBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/qualifiedbot.html)
- [QuillBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/quillbot.html)
- [Qwantbot-news — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/qwantbot-news.html)
- [Qwantbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/qwantbot.html)
- [Reflectionbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/reflectionbot.html)
- [rogerbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/rogerbot.html)
- [SBIntuitionsBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/sbintuitionsbot.html)
- [Scrapy — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/scrapy.html)
- [Screaming Frog SEO Spider — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/screaming-frog-seo-spider.html)
- [SemrushBot-BA — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/semrushbot-ba.html)
- [SemrushBot-ESI — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/semrushbot-esi.html)
- [SemrushBot-FT — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/semrushbot-ft.html)
- [SemrushBot-OCOB — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/semrushbot-ocob.html)
- [SemrushBot-SI — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/semrushbot-si.html)
- [SemrushBot-SWA — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/semrushbot-swa.html)
- [SemrushBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/semrushbot.html)
- [SEOkicks — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/seokicks.html)
- [serpstatbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/serpstatbot.html)
- [SeznamBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/seznambot.html)
- [ShapBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/shapbot.html)
- [Sidetrade indexer bot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/sidetrade-indexer-bot.html)
- [SiteAuditBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/siteauditbot.html)
- [Slackbot-LinkExpanding — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/slackbot-linkexpanding.html)
- [Slackbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/slackbot.html)
- [SplitSignalBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/splitsignalbot.html)
- [Storebot-Google — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/storebot-google.html)
- [TerraCotta — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/terracotta.html)
- [Thinkbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/thinkbot.html)
- [TikTokSpider — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/tiktokspider.html)
- [Timpibot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/timpibot.html)
- [VelenPublicWebCrawler — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/velenpublicwebcrawler.html)
- [Webzio-Extended — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/webzio-extended.html)
- [wpbot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/wpbot.html)
- [YaK — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yak.html)
- [YandexAdditional — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexadditional.html)
- [YandexAdditionalBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexadditionalbot.html)
- [YandexBlogs — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexblogs.html)
- [YandexBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexbot.html)
- [YandexCalendar — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexcalendar.html)
- [YandexComBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexcombot.html)
- [YandexDirect — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexdirect.html)
- [YandexFavicons — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexfavicons.html)
- [YandexImages — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandeximages.html)
- [YandexMarket — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexmarket.html)
- [YandexMedia — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexmedia.html)
- [YandexMetrika — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexmetrika.html)
- [YandexMobileBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexmobilebot.html)
- [YandexRenderResourcesBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexrenderresourcesbot.html)
- [YandexScreenshotBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexscreenshotbot.html)
- [YandexVideo — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexvideo.html)
- [YandexWebmaster — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yandexwebmaster.html)
- [Yeti — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/yeti.html)
- [YouBot — robots.txt token, user-agent and IP ranges](https://www.pathwren.workers.dev/crawler/youbot.html)

## data

- [Bulk data — AI Crawler Index](https://www.pathwren.workers.dev/data/index.html)

## git

- [Git repositories — AI Crawler Index](https://www.pathwren.workers.dev/git/index.html)

## ip-ranges

- [AhrefsBot and AhrefsSiteAudit IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/ahrefs-crawler.html)
- [Applebot IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/apple-applebot.html)
- [bingbot IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/bing-bingbot.html)
- [CCBot IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/commoncrawl-ccbot.html)
- [DuckDuckBot IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/duckduckgo-duckduckbot.html)
- [Googlebot IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/google-googlebot.html)
- [Google special-purpose crawlers IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/google-special.html)
- [Google-Agent (agents on Google infrastructure) IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/google-user-triggered-agents.html)
- [Google user-triggered fetchers (Google-owned ranges) IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/google-user-triggered-google.html)
- [Google user-triggered fetchers IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/google-user-triggered.html)
- [Published crawler IP ranges, one schema — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/index.html)
- [ChatGPT-User IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/openai-chatgpt-user.html)
- [GPTBot IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/openai-gptbot.html)
- [OAI-SearchBot IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/openai-searchbot.html)
- [PerplexityBot IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/perplexity-bot.html)
- [Perplexity-User IP ranges — AI Crawler Index](https://www.pathwren.workers.dev/ip-ranges/perplexity-user.html)

## operator

- [Ahrefs crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/ahrefs.html)
- [Allen Institute for AI crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/ai2.html)
- [aiHit crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/aihit.html)
- [Amazon crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/amazon.html)
- [Andi crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/andi.html)
- [Anthropic crawlers (5) — AI Crawler Index](https://www.pathwren.workers.dev/operator/anthropic.html)
- [Apple crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/apple.html)
- [Atlassian crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/atlassian.html)
- [Awario crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/awario.html)
- [Babbar crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/babbar.html)
- [Baidu crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/baidu.html)
- [ByteDance crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/bytedance.html)
- [Ceramic AI crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/ceramic.html)
- [Cloudflare crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/cloudflare.html)
- [Cohere crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/cohere.html)
- [Common Crawl crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/commoncrawl.html)
- [Crawl4AI project crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/crawl4ai.html)
- [Crawlspace crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/crawlspace.html)
- [DataForSEO crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/dataforseo.html)
- [Diffbot crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/diffbot.html)
- [Direqt crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/direqt.html)
- [DuckDuckGo crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/duckduckgo.html)
- [Echobox crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/echobox.html)
- [Exa crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/exa.html)
- [FactSet crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/factset.html)
- [Firecrawl crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/firecrawl.html)
- [Google crawlers (26) — AI Crawler Index](https://www.pathwren.workers.dev/operator/google.html)
- [Hive AI crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/hive.html)
- [Huawei crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/huawei.html)
- [Hunter (Velen) crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/hunter.html)
- [All 74 operators — AI Crawler Index](https://www.pathwren.workers.dev/operator/index.html)
- [Internet Archive crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/internetarchive.html)
- [ISS Corporate Solutions crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/iss.html)
- [Kagi crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/kagi.html)
- [Klaviyo crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/klaviyo.html)
- [LAION / img2dataset crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/laion.html)
- [Lightpanda crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/lightpanda.html)
- [Linguee crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/linguee.html)
- [Lyrenth crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/lyrenth.html)
- [Majestic crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/majestic.html)
- [Meltwater crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/meltwater.html)
- [Meta crawlers (5) — AI Crawler Index](https://www.pathwren.workers.dev/operator/meta.html)
- [Microsoft crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/microsoft.html)
- [Mistral AI crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/mistral.html)
- [Mojeek crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/mojeek.html)
- [Moz crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/moz.html)
- [Naver crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/naver.html)
- [NICT crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/nict.html)
- [OpenAI crawlers (4) — AI Crawler Index](https://www.pathwren.workers.dev/operator/openai.html)
- [Panscient crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/panscient.html)
- [Parallel crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/parallel.html)
- [Perplexity crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/perplexity.html)
- [Phind crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/phind.html)
- [Pinterest crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/pinterest.html)
- [Poseidon Research crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/poseidon.html)
- [Qualified crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/qualified.html)
- [QuantumCloud crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/quantumcloud.html)
- [QuillBot crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/quillbot.html)
- [Qwant crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/qwant.html)
- [Reflection AI crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/reflection.html)
- [ROIS-DS crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/rois.html)
- [SB Intuitions crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/sbintuitions.html)
- [Scrapy project crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/scrapy.html)
- [Screaming Frog crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/screamingfrog.html)
- [Semrush crawlers (9) — AI Crawler Index](https://www.pathwren.workers.dev/operator/semrush.html)
- [SEOkicks crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/seokicks.html)
- [Serpstat crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/serpstat.html)
- [Seznam crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/seznam.html)
- [Sidetrade crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/sidetrade.html)
- [Slack crawlers (2) — AI Crawler Index](https://www.pathwren.workers.dev/operator/slack.html)
- [Thinkbot crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/thinkbot.html)
- [Timpi crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/timpi.html)
- [Webz.io crawlers (3) — AI Crawler Index](https://www.pathwren.workers.dev/operator/webz.html)
- [Yandex crawlers (17) — AI Crawler Index](https://www.pathwren.workers.dev/operator/yandex.html)
- [You.com crawlers (1) — AI Crawler Index](https://www.pathwren.workers.dev/operator/you.html)

## policy

- [robots.txt: Allow AI search and user fetches, block the rest — AI Crawler Index](https://www.pathwren.workers.dev/policy/allow-ai-search-only.html)
- [robots.txt: Allow everything, explicitly — AI Crawler Index](https://www.pathwren.workers.dev/policy/allow-all.html)
- [robots.txt: Block AI training, keep AI search — AI Crawler Index](https://www.pathwren.workers.dev/policy/block-ai-training.html)
- [robots.txt: Block every AI crawler — AI Crawler Index](https://www.pathwren.workers.dev/policy/block-all-ai.html)
- [robots.txt: Block corpus and dataset builders — AI Crawler Index](https://www.pathwren.workers.dev/policy/block-datasets.html)
- [robots.txt: Block the crawlers with disputed robots compliance — AI Crawler Index](https://www.pathwren.workers.dev/policy/block-disputed.html)
- [robots.txt: Block SEO and backlink crawlers — AI Crawler Index](https://www.pathwren.workers.dev/policy/block-seo-tools.html)
- [Ready-made robots.txt files — AI Crawler Index](https://www.pathwren.workers.dev/policy/index.html)
- [robots.txt: Maximum AI visibility — AI Crawler Index](https://www.pathwren.workers.dev/policy/maximum-ai-visibility.html)

## snippet

- [Apache: .htaccess block — AI Crawler Index](https://www.pathwren.workers.dev/snippet/apache-htaccess.html)
- [Caddy: block by user-agent — AI Crawler Index](https://www.pathwren.workers.dev/snippet/caddy.html)
- [Cloudflare Worker: classify at the edge — AI Crawler Index](https://www.pathwren.workers.dev/snippet/cloudflare-worker.html)
- [Config snippets — AI Crawler Index](https://www.pathwren.workers.dev/snippet/index.html)
- [nginx: classify and block by user-agent — AI Crawler Index](https://www.pathwren.workers.dev/snippet/nginx-map.html)
- [Python: classify a request — AI Crawler Index](https://www.pathwren.workers.dev/snippet/python-classify.html)

## Observed clients (a second builder writes these)

- [975 client dossiers](https://www.pathwren.workers.dev/bot/index.md) — one per user-agent this host has actually served, each with the same four markdown addresses. The index links every one of them as `.md`.
- [/bot/index.html](https://www.pathwren.workers.dev/bot/index.html) — the same list as a page.
