{"csv":"clean_bot_name_full_urls.csv","totalBots":67,"counts":{"covered":48,"partial":3,"rdns":11,"ratelimit":5,"unaccounted":0},"covered":[{"bot":"anthropic-ai","category":"AI Bots","service":"anthropic","note":"Deprecated 2023-era alias; same infrastructure."},{"bot":"ChatGPT Operator","category":"AI Bots","service":"openai","note":"Retired 2025-08-31, folded into ChatGPT agent; covered by chatgpt-user.json."},{"bot":"ChatGPT-User","category":"AI Bots","service":"openai","note":null},{"bot":"ChatGPT-User/2.0","category":"AI Bots","service":"openai","note":null},{"bot":"Claude-SearchBot","category":"AI Bots","service":"anthropic","note":null},{"bot":"Claude-User","category":"AI Bots","service":"anthropic","note":null},{"bot":"Claude-Web","category":"AI Bots","service":"anthropic","note":"Deprecated 2023-era alias; same infrastructure."},{"bot":"ClaudeBot","category":"AI Bots","service":"anthropic","note":null},{"bot":"DuckAssistBot","category":"AI Bots","service":"duckduckgo","note":"duckassistbot.json is byte-identical to duckduckbot.json."},{"bot":"Google-Extended","category":"AI Bots","service":"google","note":"robots.txt control token; ranges in common-crawlers.json"},{"bot":"GPTBot","category":"AI Bots","service":"openai","note":null},{"bot":"LINER Bot","category":"AI Bots","service":"liner","note":null},{"bot":"OAI-SearchBot","category":"AI Bots","service":"openai","note":null},{"bot":"Perplexity-User","category":"AI Bots","service":"perplexity","note":null},{"bot":"Perplexity-User/1.0","category":"AI Bots","service":"perplexity","note":null},{"bot":"PerplexityBot","category":"AI Bots","service":"perplexity","note":null},{"bot":"YouBot","category":"AI Bots","service":"youbot","note":"You.com's search and answer-engine crawler, centrally operated. The docs at you.com/docs/youbot state that legitimate YouBot requests originate from 68.67.112.0/24; the youbot service scrapes that page because no JSON feed exists. Pair the CIDR with forward-confirmed rDNS on .search.you.com (PTRs look like youbot-a-b-c-d.search.you.com) in case You.com expands ranges without updating the page."},{"bot":"Amazonbot","category":"Cloud Services","service":"amazon","note":null},{"bot":"Amzn-User","category":"Cloud Services","service":"amazon","note":null},{"bot":"Google-CloudVertexBot","category":"Cloud Services","service":"google","note":null},{"bot":"Google-InspectionTool","category":"Google Bots","service":"google","note":null},{"bot":"Googlebot-Image","category":"Google Bots","service":"google","note":null},{"bot":"Googlebot-News","category":"Google Bots","service":"google","note":null},{"bot":"Googlebot-Video","category":"Google Bots","service":"google","note":null},{"bot":"GoogleOther","category":"Google Bots","service":"google","note":null},{"bot":"GoogleOther-Image","category":"Google Bots","service":"google","note":null},{"bot":"GoogleOther-Video","category":"Google Bots","service":"google","note":null},{"bot":"Storebot-Google","category":"Google Bots","service":"google","note":null},{"bot":"ia_archiver-web.archive.org","category":"Other Agents","service":"internetarchive","note":"The Wayback Machine's own capture fetcher (liveweb / Save Page Now path), distinct from the legacy Alexa ia_archiver string. IA publishes no bot JSON, rDNS scheme or Web Bot Auth key, so the registry-derived coverage is the only sound basis and it is genuine: AS7941 announces exactly 207.241.224.0/20, 208.70.24.0/21 and 2620:0:9c0::/48 per ARIN RDAP. The UA is trivially spoofable, so the ASN allowlist is the only real check."},{"bot":"TurnitinBot","category":"Other Agents","service":"turnitin","note":null},{"bot":"Amzn-SearchBot","category":"Search Engines","service":"amazon","note":null},{"bot":"Bingbot","category":"Search Engines","service":"bing","note":null},{"bot":"BingPreview","category":"Search Engines","service":"bing","note":"Microsoft's page-snapshot fetcher, folded into the evergreen bingbot agent; no BingPreview-specific docs or JSON remain (bingpreview.json 404s). Residual fetches come from the same crawl fleet, so bingbot.json is the correct and only official coverage. It is IPv4-only and rarely refreshed, so forward-confirmed rDNS to *.search.msn.com stays authoritative. Near-zero legitimate volume is the expected baseline."},{"bot":"DuckDuckBot","category":"Search Engines","service":"duckduckgo","note":null},{"bot":"Googlebot","category":"Search Engines","service":"google","note":"IMPORTANT - what the Google whitelist does NOT include: gstatic.com/ipranges/goog.json (all Google-owned IP space, ~105 prefixes) is deliberately excluded. It would admit every Google egress, including user-triggered fetch proxies anyone can drive (Apps Script UrlFetchApp, Sheets IMPORTXML, Translate proxy), turning the whitelist into a rate-limiter bypass. Coverage here is Google's five official crawler files only (~2,150 CIDRs); PSI/Lighthouse traffic is consequently NOT whitelisted - verify that via rDNS .google.com/.googlebot.com."},{"bot":"MicrosoftPreview","category":"Search Engines","service":"bing","note":"Link-preview and page-snapshot fetcher for Microsoft products (Teams, Outlook, Office, Copilot). The self-identifying link in its own UA, aka.ms/MicrosoftPreview, redirects to Bing's crawler doc, which lists MicrosoftPreview alongside Bingbot and routes it to the same verification, so bingbot.json applies. Authoritative fallback is forward-confirmed rDNS to *.search.msn.com."},{"bot":"Mojeek","category":"Search Engines","service":"mojeek","note":null},{"bot":"MojeekBot","category":"Search Engines","service":"mojeek","note":null},{"bot":"AhrefsBot","category":"SEO Tools","service":"ahrefs","note":null},{"bot":"facebookexternalhit","category":"Social Bots","service":"meta","note":"Via AS32934 BGP prefixes (Meta publishes no list; ASN is their historical guidance)."},{"bot":"Meta-ExternalAds","category":"Social Bots","service":"meta","note":null},{"bot":"meta-externalads","category":"Social Bots","service":"meta","note":null},{"bot":"Meta-ExternalAgent","category":"Social Bots","service":"meta","note":null},{"bot":"meta-externalagent","category":"Social Bots","service":"meta","note":null},{"bot":"Meta-ExternalFetcher","category":"Social Bots","service":"meta","note":null},{"bot":"meta-externalfetcher","category":"Social Bots","service":"meta","note":null},{"bot":"Meta-WebIndexer","category":"Social Bots","service":"meta","note":null},{"bot":"meta-webindexer","category":"Social Bots","service":"meta","note":null}],"partial":[{"bot":"archive.org_bot","category":"Other Agents","service":"internetarchive","note":"The Internet Archive's archival crawler UA (Heritrix/Brozzler plus Archive-It subscriber crawls). IA publishes no IP list and hands allowlist data out privately via Archive-It support, so the ARIN AS7941 coverage catches the bulk but misses IPv6 2620:0:9c0::/48 and Internet Archive Canada (AS399784: 204.62.246.0/23, 204.62.248.0/23). rDNS is useless here - most crawl IPs are NXDOMAIN - and the UA is trivially spoofed."},{"bot":"FacebookBot","category":"Social Bots","service":"meta","note":"Meta's legacy crawler for harvesting public web text to train language and speech models, superseded by meta-externalagent in 2024 and no longer documented on Meta's developer site. Meta publishes no IP file for any crawler, so BGP prefixes originated by AS32934 (878 announced) are the only workable allowlist - coarse, since they admit every other Meta service. A FacebookBot UA from outside AS32934 is almost certainly spoofed."},{"bot":"Twitterbot","category":"Social Bots","service":"x-twitter","note":"X's link-unfurling crawler, fetching URLs posted in tweets to build Card previews. X publishes nothing machine-readable anymore, so the three hardcoded CIDRs are frozen snapshots; all three verify in ARIN RDAP as direct allocations to Twitter Inc. (AS13414), but X announces further AS13414 prefixes outside the list. Treat as high-precision but incomplete and fall back to an AS13414 check - rDNS is unreliable here."}],"rdns":[{"bot":"Gemini-Deep-Research","category":"Google Bots","service":"google","note":"UA seen when Gemini's Deep Research mode fans out to fetch pages for a user's prompt; user-triggered, not a corpus crawler. Google documents no such token, so no per-bot IP file exists - real traffic leaves Google-owned space and will match the user-triggered fetcher/agent files or goog.json, but that mapping is inferred. Harden with rDNS to google.com/googlebot.com or Web Bot Auth against agent.bot.goog."},{"bot":"adidxbot","category":"Search Engines","service":"bing","note":"Microsoft Advertising's ad-side crawler: it fetches ad landing pages and feed assets for policy and quality checks, so blocking it causes ad disapprovals. No adidxbot-specific file exists (adidxbot.json 404s) and bingbot.json is labelled for Bingbot only, though live PTRs across its ranges all resolve to msnbot-*.search.msn.com. Keep forward-confirmed rDNS to *.search.msn.com as the authoritative check."},{"bot":"Applebot","category":"Search Engines","service":"apple","note":"Apple's search/AI crawler feeding Spotlight, Siri and Safari search. applebot.json is the only official source but is stale (creationTime 2023-10-27, 12 IPv4 CIDRs, no IPv6) and provably incomplete - Apple's own documented example host 17.58.101.179 sits outside every prefix. Treat the JSON as a fast-path allow and make forward-confirmed rDNS on *.applebot.apple.com authoritative; all traffic is inside 17.0.0.0/8 (AS714)."},{"bot":"BingVideoPreview","category":"Search Engines","service":"bing","note":"Microsoft's video-preview fetcher, pulling video files and thumbnails for hover-play previews on the same fleet as bingbot. bingbot.json is live but carries no user-agent metadata and no per-agent list exists, so nothing from Microsoft says video fetches stay inside it. Use those ranges as the primary allow, not a hard block, and verify with rDNS to *.search.msn.com (AS8075/AS8068)."},{"bot":"Slurp","category":"Search Engines","service":null,"note":"Yahoo's own search crawler, still live - the crawl.yahoo.net zone is actively maintained. Yahoo has never published a crawler IP file and its help articles document only the UA, so the sole defensible check is forward-confirmed rDNS ending in .crawl.yahoo.net. Do not use ASN: crawl.yahoo.net now fronts on AWS space alongside legacy Yahoo ranges, and the old published IPs are stale."},{"bot":"Google Page Speed Insights","category":"SEO Tools","service":null,"note":"PageSpeed Insights runs Lighthouse server-side in rotating Google datacenters, so this is a real Google fleet, but Google staff state on the record that no fixed IP set will be published and PSI appears in none of the five official crawler IP files. Verify by forward-confirmed rDNS ending in .google.com or .googlebot.com; gstatic.com/ipranges/goog.json exists but whitelisting it would admit all Google egress incl. user-triggered fetch proxies (Apps Script etc.) - a rate-limiter bypass."},{"bot":"SemrushBot","category":"SEO Tools","service":null,"note":"Semrush's backlink/SEO index crawler. Semrush states verbatim that it does not use consecutive IP blocks and that operators should not block it by IP; the only documented CIDRs cover SiteAuditBot and SemrushBot-SI, not generic SemrushBot. Verify with forward-confirmed rDNS on the PTR suffix .bot.semrush.com (never a bare semrush.com substring), backstopped by AS209366."},{"bot":"SemrushBot-OCOB","category":"SEO Tools","service":null,"note":"Semrush's Content Toolkit crawler for the AI article generator and content optimizer. Semrush names it on semrush.com/bot but publishes ranges for only two of its ten crawlers, neither of which covers OCOB, and there is no bot JSON or Web Bot Auth key. Verify with forward-confirmed rDNS under *.bot.semrush.com, optionally constrained to AS209366; treat an OCOB UA failing that as spoofed."},{"bot":"SemrushBotSwa","category":"SEO Tools","service":null,"note":"SemrushBot-SWA is the SEO Writing Assistant URL fetcher, triggered when a user pastes a URL and running on Semrush's own fleet. It is named only as a robots.txt token - the single KB article with IPs covers SemrushBot-SI and SiteAuditBot, not SWA - and Semrush advises against IP blocking outright. Best available signals are the rDNS suffix .bot.semrush.com and origin AS209366, both heuristics."},{"bot":"LinkedInBot","category":"Social Bots","service":null,"note":"LinkedIn's centrally operated Open Graph link-preview fetcher, still active, but LinkedIn publishes no CIDR list, no bot JSON and no Web Bot Auth directory - robots.txt and whitelist-crawl@linkedin.com are the only official channels. Verify with forward-confirmed rDNS ending in .linkedin.com (PTRs look like 108-174-2-1.fwd.linkedin.com), with AS14413 as a coarse secondary signal."},{"bot":"Pinterestbot","category":"Social Bots","service":"pinterest","note":"Pinterest's content crawler, fetching pages users pin plus Rich Pin metadata. The only official IP statement covers US crawls from 54.236.1.0/24 - live PTRs across that /24 resolve to crawl-*.pinterestcrawler.com and crawl-*.pinterest.com, so the range is genuine and fully crawler-allocated - but non-US crawls have no fixed range. Pair it with forward-confirmed rDNS on pinterest.com / pinterestcrawler.com."}],"ratelimit":[{"bot":"Grok","category":"AI Bots","note":"Live-retrieval fetcher behind xAI's Grok assistant (sibling tokens xAI-Bot, xAI-Grok, Grok-DeepSearch). xAI publishes no crawler docs, no IP JSON and no Web Bot Auth directory, and operator reports show the real fetcher often hides behind stale browser UAs on rotating hosting/residential proxies (M247, Datacamp), so no CIDR or rDNS suffix holds. Rate-limit rather than allowlist."},{"bot":"GrokBot","category":"AI Bots","note":"Community-propagated token for xAI's Grok crawler fleet; xAI has never officially attested it - its own robots.txt names GPTBot, ClaudeBot and Applebot-Extended but no xAI token. No IP file, no rDNS convention and no Web Bot Auth directory exist, and traffic reportedly arrives from rotating proxy ranges under spoofed browser UAs. Unverifiable today; re-check on future sweeps."},{"bot":"peer39_crawler/1.0","category":"Other Agents","note":"Peer39's contextual ad-verification crawler - a genuine centrally operated fleet worth allowing, but the company publishes no IP list, no crawler JSON and no Web Bot Auth entry, and its address space serves no PTR records at all. The only non-UA signal is ASN ownership: AS54183 (Peer39 Inc) originates 15 IPv4 prefixes including 204.15.208.0/22. Treat an ASN match as corroboration, not verification."},{"bot":"Quora-Bot","category":"Social Bots","note":"A real, low-volume server-side fetcher Quora runs to build link previews for URLs shared in answers, so it is a legitimate whitelisting candidate. Quora publishes no crawler docs, no bot JSON, no rDNS convention and no Web Bot Auth directory; robotstxt@quora.com in its robots.txt is the only official channel. UA-only for now - keep on the radar."},{"bot":"Slackbot","category":"Social Bots","note":"Slack's centrally operated fetcher family (Slackbot, Slackbot-LinkExpanding, Slack-ImgProxy) publishes no IP source: api.slack.com/robots documents the UA strings only, egress is rotating AWS with plain ec2-* rDNS, and there is no Web Bot Auth directory. Match the UA prefix for crawl traffic; for inbound Slack webhooks verify the X-Slack-Signature HMAC instead of any IP rule."}],"unaccounted":[]}