# AAN Botpass

official crawler IP ranges, polled daily, compiled for the rate limiter

Every bot in clean_bot_name_full_urls.csv (67 bots) cross-referenced against the running services. Bots added to the CSV without a verdict appear under unaccounted — that's the research queue.

# covered 48

A service polls an official source for these bots.

BotCategoryServiceNote
anthropic-ai AI Bots anthropic Deprecated 2023-era alias; same infrastructure.
ChatGPT Operator AI Bots openai Retired 2025-08-31, folded into ChatGPT agent; covered by chatgpt-user.json.
ChatGPT-User AI Bots openai
ChatGPT-User/2.0 AI Bots openai
Claude-SearchBot AI Bots anthropic
Claude-User AI Bots anthropic
Claude-Web AI Bots anthropic Deprecated 2023-era alias; same infrastructure.
ClaudeBot AI Bots anthropic
DuckAssistBot AI Bots duckduckgo duckassistbot.json is byte-identical to duckduckbot.json.
Google-Extended AI Bots google robots.txt control token; ranges in common-crawlers.json
GPTBot AI Bots openai
LINER Bot AI Bots liner
OAI-SearchBot AI Bots openai
Perplexity-User AI Bots perplexity
Perplexity-User/1.0 AI Bots perplexity
PerplexityBot AI Bots perplexity
YouBot AI Bots youbot You.com's search and answer-engine crawler, centrally operated. The docs at you.com/docs/youbot state that legitimate YouBot requests originate from 68.67.112.0/24; the youbot service scrapes that page because no JSON feed exists. Pair the CIDR with forward-confirmed rDNS on .search.you.com (PTRs look like youbot-a-b-c-d.search.you.com) in case You.com expands ranges without updating the page.
Amazonbot Cloud Services amazon
Amzn-User Cloud Services amazon
Google-CloudVertexBot Cloud Services google
Google-InspectionTool Google Bots google
Googlebot-Image Google Bots google
Googlebot-News Google Bots google
Googlebot-Video Google Bots google
GoogleOther Google Bots google
GoogleOther-Image Google Bots google
GoogleOther-Video Google Bots google
Storebot-Google Google Bots google
ia_archiver-web.archive.org Other Agents internetarchive The Wayback Machine's own capture fetcher (liveweb / Save Page Now path), distinct from the legacy Alexa ia_archiver string. IA publishes no bot JSON, rDNS scheme or Web Bot Auth key, so the registry-derived coverage is the only sound basis and it is genuine: AS7941 announces exactly 207.241.224.0/20, 208.70.24.0/21 and 2620:0:9c0::/48 per ARIN RDAP. The UA is trivially spoofable, so the ASN allowlist is the only real check.
TurnitinBot Other Agents turnitin
Amzn-SearchBot Search Engines amazon
Bingbot Search Engines bing
BingPreview Search Engines bing Microsoft's page-snapshot fetcher, folded into the evergreen bingbot agent; no BingPreview-specific docs or JSON remain (bingpreview.json 404s). Residual fetches come from the same crawl fleet, so bingbot.json is the correct and only official coverage. It is IPv4-only and rarely refreshed, so forward-confirmed rDNS to *.search.msn.com stays authoritative. Near-zero legitimate volume is the expected baseline.
DuckDuckBot Search Engines duckduckgo
Googlebot Search Engines google IMPORTANT - what the Google whitelist does NOT include: gstatic.com/ipranges/goog.json (all Google-owned IP space, ~105 prefixes) is deliberately excluded. It would admit every Google egress, including user-triggered fetch proxies anyone can drive (Apps Script UrlFetchApp, Sheets IMPORTXML, Translate proxy), turning the whitelist into a rate-limiter bypass. Coverage here is Google's five official crawler files only (~2,150 CIDRs); PSI/Lighthouse traffic is consequently NOT whitelisted - verify that via rDNS .google.com/.googlebot.com.
MicrosoftPreview Search Engines bing Link-preview and page-snapshot fetcher for Microsoft products (Teams, Outlook, Office, Copilot). The self-identifying link in its own UA, aka.ms/MicrosoftPreview, redirects to Bing's crawler doc, which lists MicrosoftPreview alongside Bingbot and routes it to the same verification, so bingbot.json applies. Authoritative fallback is forward-confirmed rDNS to *.search.msn.com.
Mojeek Search Engines mojeek
MojeekBot Search Engines mojeek
AhrefsBot SEO Tools ahrefs
facebookexternalhit Social Bots meta Via AS32934 BGP prefixes (Meta publishes no list; ASN is their historical guidance).
Meta-ExternalAds Social Bots meta
meta-externalads Social Bots meta
Meta-ExternalAgent Social Bots meta
meta-externalagent Social Bots meta
Meta-ExternalFetcher Social Bots meta
meta-externalfetcher Social Bots meta
Meta-WebIndexer Social Bots meta
meta-webindexer Social Bots meta

# partial 3

Real IP coverage, but the caveat matters operationally.

BotCategoryServiceNote
archive.org_bot Other Agents internetarchive The Internet Archive's archival crawler UA (Heritrix/Brozzler plus Archive-It subscriber crawls). IA publishes no IP list and hands allowlist data out privately via Archive-It support, so the ARIN AS7941 coverage catches the bulk but misses IPv6 2620:0:9c0::/48 and Internet Archive Canada (AS399784: 204.62.246.0/23, 204.62.248.0/23). rDNS is useless here - most crawl IPs are NXDOMAIN - and the UA is trivially spoofed.
FacebookBot Social Bots meta Meta's legacy crawler for harvesting public web text to train language and speech models, superseded by meta-externalagent in 2024 and no longer documented on Meta's developer site. Meta publishes no IP file for any crawler, so BGP prefixes originated by AS32934 (878 announced) are the only workable allowlist - coarse, since they admit every other Meta service. A FacebookBot UA from outside AS32934 is almost certainly spoofed.
Twitterbot Social Bots x-twitter X's link-unfurling crawler, fetching URLs posted in tweets to build Card previews. X publishes nothing machine-readable anymore, so the three hardcoded CIDRs are frozen snapshots; all three verify in ARIN RDAP as direct allocations to Twitter Inc. (AS13414), but X announces further AS13414 prefixes outside the list. Treat as high-precision but incomplete and fall back to an AS13414 check - rDNS is unreliable here.

# rdns 11

The operator’s authoritative check is forward-confirmed reverse DNS; IP coverage is partial or absent. The note names the required PTR suffix.

BotCategoryServiceNote
Gemini-Deep-Research Google Bots google UA seen when Gemini's Deep Research mode fans out to fetch pages for a user's prompt; user-triggered, not a corpus crawler. Google documents no such token, so no per-bot IP file exists - real traffic leaves Google-owned space and will match the user-triggered fetcher/agent files or goog.json, but that mapping is inferred. Harden with rDNS to google.com/googlebot.com or Web Bot Auth against agent.bot.goog.
adidxbot Search Engines bing Microsoft Advertising's ad-side crawler: it fetches ad landing pages and feed assets for policy and quality checks, so blocking it causes ad disapprovals. No adidxbot-specific file exists (adidxbot.json 404s) and bingbot.json is labelled for Bingbot only, though live PTRs across its ranges all resolve to msnbot-*.search.msn.com. Keep forward-confirmed rDNS to *.search.msn.com as the authoritative check.
Applebot Search Engines apple Apple's search/AI crawler feeding Spotlight, Siri and Safari search. applebot.json is the only official source but is stale (creationTime 2023-10-27, 12 IPv4 CIDRs, no IPv6) and provably incomplete - Apple's own documented example host 17.58.101.179 sits outside every prefix. Treat the JSON as a fast-path allow and make forward-confirmed rDNS on *.applebot.apple.com authoritative; all traffic is inside 17.0.0.0/8 (AS714).
BingVideoPreview Search Engines bing Microsoft's video-preview fetcher, pulling video files and thumbnails for hover-play previews on the same fleet as bingbot. bingbot.json is live but carries no user-agent metadata and no per-agent list exists, so nothing from Microsoft says video fetches stay inside it. Use those ranges as the primary allow, not a hard block, and verify with rDNS to *.search.msn.com (AS8075/AS8068).
Slurp Search Engines Yahoo's own search crawler, still live - the crawl.yahoo.net zone is actively maintained. Yahoo has never published a crawler IP file and its help articles document only the UA, so the sole defensible check is forward-confirmed rDNS ending in .crawl.yahoo.net. Do not use ASN: crawl.yahoo.net now fronts on AWS space alongside legacy Yahoo ranges, and the old published IPs are stale.
Google Page Speed Insights SEO Tools PageSpeed Insights runs Lighthouse server-side in rotating Google datacenters, so this is a real Google fleet, but Google staff state on the record that no fixed IP set will be published and PSI appears in none of the five official crawler IP files. Verify by forward-confirmed rDNS ending in .google.com or .googlebot.com; gstatic.com/ipranges/goog.json exists but whitelisting it would admit all Google egress incl. user-triggered fetch proxies (Apps Script etc.) - a rate-limiter bypass.
SemrushBot SEO Tools Semrush's backlink/SEO index crawler. Semrush states verbatim that it does not use consecutive IP blocks and that operators should not block it by IP; the only documented CIDRs cover SiteAuditBot and SemrushBot-SI, not generic SemrushBot. Verify with forward-confirmed rDNS on the PTR suffix .bot.semrush.com (never a bare semrush.com substring), backstopped by AS209366.
SemrushBot-OCOB SEO Tools Semrush's Content Toolkit crawler for the AI article generator and content optimizer. Semrush names it on semrush.com/bot but publishes ranges for only two of its ten crawlers, neither of which covers OCOB, and there is no bot JSON or Web Bot Auth key. Verify with forward-confirmed rDNS under *.bot.semrush.com, optionally constrained to AS209366; treat an OCOB UA failing that as spoofed.
SemrushBotSwa SEO Tools SemrushBot-SWA is the SEO Writing Assistant URL fetcher, triggered when a user pastes a URL and running on Semrush's own fleet. It is named only as a robots.txt token - the single KB article with IPs covers SemrushBot-SI and SiteAuditBot, not SWA - and Semrush advises against IP blocking outright. Best available signals are the rDNS suffix .bot.semrush.com and origin AS209366, both heuristics.
LinkedInBot Social Bots LinkedIn's centrally operated Open Graph link-preview fetcher, still active, but LinkedIn publishes no CIDR list, no bot JSON and no Web Bot Auth directory - robots.txt and whitelist-crawl@linkedin.com are the only official channels. Verify with forward-confirmed rDNS ending in .linkedin.com (PTRs look like 108-174-2-1.fwd.linkedin.com), with AS14413 as a coarse secondary signal.
Pinterestbot Social Bots pinterest Pinterest's content crawler, fetching pages users pin plus Rich Pin metadata. The only official IP statement covers US crawls from 54.236.1.0/24 - live PTRs across that /24 resolve to crawl-*.pinterestcrawler.com and crawl-*.pinterest.com, so the range is genuine and fully crawler-allocated - but non-US crawls have no fixed range. Pair it with forward-confirmed rDNS on pinterest.com / pinterestcrawler.com.

# ratelimit 5

No ranges, no rDNS scheme — unverifiable. Rate-limit these user agents instead of whitelisting.

BotCategoryWhy / best signal
Grok AI Bots Live-retrieval fetcher behind xAI's Grok assistant (sibling tokens xAI-Bot, xAI-Grok, Grok-DeepSearch). xAI publishes no crawler docs, no IP JSON and no Web Bot Auth directory, and operator reports show the real fetcher often hides behind stale browser UAs on rotating hosting/residential proxies (M247, Datacamp), so no CIDR or rDNS suffix holds. Rate-limit rather than allowlist.
GrokBot AI Bots Community-propagated token for xAI's Grok crawler fleet; xAI has never officially attested it - its own robots.txt names GPTBot, ClaudeBot and Applebot-Extended but no xAI token. No IP file, no rDNS convention and no Web Bot Auth directory exist, and traffic reportedly arrives from rotating proxy ranges under spoofed browser UAs. Unverifiable today; re-check on future sweeps.
peer39_crawler/1.0 Other Agents Peer39's contextual ad-verification crawler - a genuine centrally operated fleet worth allowing, but the company publishes no IP list, no crawler JSON and no Web Bot Auth entry, and its address space serves no PTR records at all. The only non-UA signal is ASN ownership: AS54183 (Peer39 Inc) originates 15 IPv4 prefixes including 204.15.208.0/22. Treat an ASN match as corroboration, not verification.
Quora-Bot Social Bots A real, low-volume server-side fetcher Quora runs to build link previews for URLs shared in answers, so it is a legitimate whitelisting candidate. Quora publishes no crawler docs, no bot JSON, no rDNS convention and no Web Bot Auth directory; robotstxt@quora.com in its robots.txt is the only official channel. UA-only for now - keep on the radar.
Slackbot Social Bots Slack's centrally operated fetcher family (Slackbot, Slackbot-LinkExpanding, Slack-ImgProxy) publishes no IP source: api.slack.com/robots documents the UA strings only, egress is rotating AWS with plain ec2-* rDNS, and there is no Web Bot Auth directory. Match the UA prefix for crawl traffic; for inbound Slack webhooks verify the X-Slack-Signature HMAC instead of any IP rule.