This is the full roster of every crawler, fetcher and agent that AI Crawler Check tests your site against: 248 bots across 8 categories, operated by 103 companies and projects. Each token links to its own directory profile with the exact robots.txt rule, the operator's documentation and a safety verdict.
The list is generated from the same database the checker runs on, so it cannot drift from what the tool actually detects. It was last rebuilt for release 6.1, which checked every token against its operator's own documentation, labelled each one Official or Reported, and removed the strings that turned out to be dead spellings or bare product names. Before reading on, run a free AI crawler check so you know which of these can currently reach you.
Key Takeaways
- 88 of the 248 are AI and LLM crawlers. The rest are search engines, Google fetchers, SEO tools, social unfurlers, scrapers and cloud agents, and they need different rules.
- 211 tokens are rated Safe, 32 Caution and 5 Aggressive. Aggressive means documented evidence of ignoring robots.txt or excessive crawl rates.
- 52 tokens are user-triggered fetchers (ChatGPT-User, Claude-User, Google-GeminiNotebook, Perplexity-User, Kimi-User). They fetch a page because a person asked for it and, by their operators' own statements, may not honour robots.txt.
- 177 are named on their operator's own documentation and 71 are Reported: named by a credible public registry but not by the operator. Every profile shows which, and links the source, so you can decide how much weight to give the rule.
The Shape of the Database
Not every bot is an AI bot, and treating them all the same is the most common mistake we see in robots.txt files. A rule that blocks Googlebot because it sat in the same block as GPTBot costs you organic rankings for nothing. The eight categories below exist so that you can write one policy per class of traffic rather than one per token.
Counts are generated from the live database at build time and sum to 248.
By purpose, 28 tokens exist to collect training data, 52 fetch a page on a user's request (these usually ignore robots.txt by design, because a human asked for the page), 67 build a search or answer index, 11 do both training and search, and the remainder scrape, monitor or unfurl links. The distinction matters because the on-demand fetchers are the ones that put your page in front of a person in real time; blocking them removes you from the answer while the training crawlers keep reading. Our guide to training bots vs search bots walks through the policy implications.
New in Release 6.1: 155 Additions
Additions come in two tiers, and every profile wears its tier as a badge. Official docs (177 of them) means the operator's own crawler page names the token and the profile links to it. Reported (71 of them) means no operator page exists but a credible public source does (the ai.robots.txt registry, Trakkr's observatory, user-agents.net); the profile says so plainly and links that source. Nothing is listed on hearsay alone, and bare product names such as Grok, OpenAI or Operator are not tokens and are not listed.
Google documented fetchers
- Google-GeminiNotebook (Google, on-demand fetch). Safe.
- Google-CWS (Google, monitoring). Safe.
- GoogleMessages (Google, link preview). Safe.
- Google-Pinpoint (Google, on-demand fetch). Safe.
Moonshot AI (Kimi)
- KimiBot (Moonshot AI, AI training). Caution.
- Kimi-SearchBot (Moonshot AI, indexing). Safe.
New fetchers from known operators
- Meta-ExternalAds (Meta, monitoring). Safe.
- MistralAI-Training (Mistral AI, AI training). Safe.
- Diffbot-User (Diffbot, on-demand fetch). Safe.
- ShapBot (Parallel, indexing). Caution.
- Brightbot (Bright Data, AI training). Aggressive.
Restored after the second verification pass (each is on its vendor page)
- Slurp (Yahoo, indexing). Safe.
- Bytespider (ByteDance, training + search). Aggressive.
- ExaSearchBot (Exa, indexing). Safe.
- KlaviyoAIBot (Klaviyo, AI training). Safe.
- SBIntuitionsBot (SB Intuitions, AI training). Safe.
- Arquivo-web-crawler (Arquivo.pt, scraping). Safe.
- Cotoyogi (ROIS-DS, AI training). Safe.
- Webzio-Extended (Webz.io, AI training). Caution.
Yandex robots, from the official Yandex list
- YaDirectFetcher (Yandex, monitoring). Safe.
- YandexAccessibilityBot (Yandex, indexing). Safe.
- YandexAdNet (Yandex, monitoring). Safe.
- YandexBlogs (Yandex, indexing). Safe.
- YandexCalendar (Yandex, on-demand fetch). Safe.
- YandexCheckBot (Yandex, monitoring). Safe.
- YandexComBot (Yandex, indexing). Safe.
- YandexDirect (Yandex, monitoring). Safe.
- YandexDirectDyn (Yandex, monitoring). Safe.
- YandexFavicons (Yandex, indexing). Safe.
- YandexFeed (Yandex, indexing). Safe.
- YandexImageResizer (Yandex, on-demand fetch). Safe.
- YandexImages (Yandex, indexing). Safe.
- YandexMarket (Yandex, indexing). Safe.
- YandexMedia (Yandex, indexing). Safe.
- YandexMetrika (Yandex, monitoring). Safe.
- YandexMobileBot (Yandex, indexing). Safe.
- YandexMobileScreenShotBot (Yandex, monitoring). Safe.
- YandexOntoDB (Yandex, indexing). Safe.
- YandexOntoDBAPI (Yandex, on-demand fetch). Safe.
- YandexPagechecker (Yandex, monitoring). Safe.
- YandexPartner (Yandex, monitoring). Safe.
- YandexRCA (Yandex, indexing). Safe.
- YandexRenderResourcesBot (Yandex, indexing). Safe.
- YandexScreenshotBot (Yandex, monitoring). Safe.
- YandexSearchShop (Yandex, indexing). Safe.
- YandexSitelinks (Yandex, monitoring). Safe.
- YandexSpravBot (Yandex, indexing). Safe.
- YandexTracker (Yandex, on-demand fetch). Safe.
- YandexUserproxy (Yandex, on-demand fetch). Safe.
- YandexVerticals (Yandex, indexing). Safe.
- YandexVertis (Yandex, indexing). Safe.
- YandexVideo (Yandex, indexing). Safe.
- YandexVideoParser (Yandex, indexing). Safe.
- YandexWebmaster (Yandex, monitoring). Safe.
Baidu product spiders
- Baiduspider-image (Baidu, indexing). Safe.
- Baiduspider-video (Baidu, indexing). Safe.
- Baiduspider-news (Baidu, indexing). Safe.
- Baiduspider-favo (Baidu, indexing). Safe.
- Baiduspider-cpro (Baidu, monitoring). Safe.
- Baiduspider-ads (Baidu, monitoring). Safe.
Google special-case crawlers
- APIs-Google (Google, on-demand fetch). Safe.
- AdsBot-Google (Google, monitoring). Safe.
- AdsBot-Google-Mobile (Google, monitoring). Safe.
- Mediapartners-Google (Google, monitoring). Safe.
- Google-Safety (Google, monitoring). Safe.
Bing crawlers
- AdIdxBot (Microsoft, monitoring). Safe.
- BingPreview (Microsoft, indexing). Safe.
- BingVideoPreview (Microsoft, indexing). Safe.
- MicrosoftPreview (Microsoft, link preview). Safe.
Search, AI and SEO operators with documented sibling tokens
- Kagibot (Kagi, indexing). Safe.
- LinerBot (LINER, training + search). Safe.
- SBIntuitions-SearchBot (SB Intuitions, on-demand fetch). Safe.
- AhrefsSiteAudit (Ahrefs, monitoring). Safe.
- Rogerbot (Moz, monitoring). Safe.
- AwarioBot (Awario, monitoring). Safe.
- Slackbot-LinkExpanding (Slack, link preview). Safe.
ChatGPT Agent, confirmed by OpenAI's allowlisting guide (a signed agent, not a robots.txt token)
- ChatGPT Agent (OpenAI, on-demand fetch). Safe.
- ChatGPT Operator (OpenAI, on-demand fetch). Safe.
Vendors found through the Trakkr cross-check whose own pages print the token
- AIWebIndex (Lyrenth, indexing). Safe.
- LinkupBot (Linkup, indexing). Safe.
- Mozilla-Tabstack (Mozilla, on-demand fetch). Safe.
- TerraCotta (Ceramic AI, AI training). Caution.
- Panscient (Panscient, scraping). Safe.
- SemrushBot-SWA (Semrush, monitoring). Safe.
Reported tokens: no operator page, but a credible public registry names them
- AddSearchBot (AddSearch, indexing). Safe.
- AI2Bot-DeepResearchEval (Allen Institute for AI, on-demand fetch). Safe.
- aiHitBot (aiHit, AI training). Caution.
- amazon-QBusiness (Amazon, indexing). Safe.
- AmazonBuyForMe (Amazon, on-demand fetch). Safe.
- Andibot (Andi, training + search). Safe.
- anthropic-ai (Anthropic, AI training). Safe.
- ApifyBot (Apify, scraping). Caution.
- atlassian-bot (Atlassian, indexing). Safe.
- AzureAI-SearchBot (Microsoft, indexing). Safe.
- BLEXBot (WebMeUp, scraping). Safe.
- Bravebot (Brave, on-demand fetch). Safe.
- ChatGLM-Spider (Zhipu AI, AI training). Caution.
- Chrome-Lighthouse (Google, monitoring). Safe.
- Claude-Code (Anthropic, on-demand fetch). Safe.
- Claude-Web (Anthropic, on-demand fetch). Safe.
- Cloudflare-AutoRAG (Cloudflare, indexing). Safe.
- cohere-ai (Cohere, training + search). Safe.
- cohere-training-data-crawler (Cohere, AI training). Caution.
- Crawlspace (Crawlspace, scraping). Safe.
- Cursor (Anysphere, on-demand fetch). Safe.
- DeepSeekBot (DeepSeek, training + search). Aggressive.
- Devin (Cognition AI, on-demand fetch). Safe.
- DoubaoBot (ByteDance, training + search). Caution.
- Echobot Bot (Dealfront, scraping). Safe.
- EchoboxBot (Echobox, scraping). Safe.
- ERNIEBot (Baidu, AI training). Caution.
- ExaBot (Exa, indexing). Safe.
- FacebookBot (Meta, link preview). Safe.
- Factset_spyderbot (FactSet, scraping). Caution.
- FriendlyCrawler (FriendlyCrawler, scraping). Safe.
- Gemini-Deep-Research (Google, on-demand fetch). Safe.
- Google-Firebase (Google, on-demand fetch). Safe.
- Google-Gemini-CLI (Google, on-demand fetch). Safe.
- Google-NotebookLM (Google, on-demand fetch). Safe.
- GoogleAgent-Mariner (Google, on-demand fetch). Safe.
- GoogleAgent-URLContext (Google, on-demand fetch). Safe.
- GrokBot (xAI, training + search). Safe.
- ia_archiver (Internet Archive, scraping). Caution.
- iAskBot (iAsk, training + search). Safe.
- iaskspider (iAsk, scraping). Safe.
- IbouBot (Ibou, indexing). Safe.
- ISSCyberRiskCrawler (ISS, scraping). Safe.
- kagi-fetcher (Kagi, on-demand fetch). Safe.
- Kangaroo Bot (Kangaroo, scraping). Safe.
- LAIONDownloader (LAION, AI training). Caution.
- Manus-User (Manus, on-demand fetch). Safe.
- MauiBot (MauiBot, scraping). Safe.
- MyCentralAIScraperBot (MyCentral AI, AI training). Safe.
- netEstate Imprint Crawler (netEstate, scraping). Safe.
- Netvibes (Dassault Systemes (Netvibes), monitoring). Safe.
- NovaAct (Amazon, scraping). Safe.
- opencode (OpenCode, on-demand fetch). Safe.
- PanguBot (Huawei, AI training). Safe.
- PhindBot (Phind, on-demand fetch). Safe.
- Poseidon Research Crawler (Poseidon Research, scraping). Safe.
- QualifiedBot (Qualified, scraping). Safe.
- QuillBot (QuillBot, on-demand fetch). Caution.
- Quora-Bot (Quora, link preview). Safe.
- QwenBot (Alibaba, AI training). Caution.
- Shap-User (Parallel, on-demand fetch). Safe.
- Sidetrade Indexer Bot (Sidetrade, monitoring). Safe.
- TavilyBot (Tavily, on-demand fetch). Safe.
- TikTokSpider (ByteDance, AI training). Aggressive.
- Timpibot (Timpi, scraping). Safe.
- TongyiBot (Alibaba, on-demand fetch). Caution.
- Trae (ByteDance, on-demand fetch). Safe.
- WRTNBot (WRTN, training + search). Safe.
- xAI-SearchBot (xAI, indexing). Caution.
- YiyanBot (Baidu, on-demand fetch). Caution.
- YouBot (You, training + search). Safe.
Official, Reported, Retired: How to Read the Badges
Most crawler lists mix three very different things: tokens the vendor publishes, tokens that only registries and server logs know, and dead spellings that match nothing. We separate them, because the right action is different for each.
Official docs (177)
The operator's own page names the token. Purpose, robots.txt behaviour and, where published, IP ranges come from that page. OpenAI, Anthropic, Google, Meta, Apple, Amazon, Microsoft, Mistral, Perplexity, Moonshot, Yandex, Baidu and ByteDance's Toutiao Search team all publish such pages.
Reported (71)
No operator page, but a credible public source names it. Claude-Web, anthropic-ai, GrokBot, xAI-SearchBot, DeepSeekBot, QwenBot, cohere-ai, Gemini-Deep-Research and the AI coding agents sit here. The rule on the profile is the best available spelling; treat it as best effort and pair it with a firewall rule if the traffic matters.
Retired or never a token
anthropic-aireplaced byClaudeBotclaude-webreplaced byClaudeBotgoogle-notebooklmreplaced byGoogle-GeminiNotebooknotebooklmreplaced byGoogle-GeminiNotebookchatgpt operatorreplaced byChatGPT-Userchatgpt agentreplaced byChatGPT-Useroperatorreplaced byChatGPT-Userteoma
The validator flags these and names the token to add. ChatGPT Agent and ChatGPT Operator are real, but they are CDN bot-manager names for OpenAI's signed Cloud browser, not robots.txt tokens; their profiles explain Web Bot Auth.
Deliberately absent: dead Yahoo spellings (YahooSeeker, Yahoo-MMCrawler, Yahoo-Blogs, Yahoo-FeedSeeker), Teoma, Ai2Bot-Dolma (AI2's page names only AI2Bot), and bare company or product names that were never user-agent tokens. The full exclusion list, with a reason for each, ships with the site's source data.
The Full Roster, by Category
Tokens are the exact strings to use in User-agent: lines. Matching is case-insensitive in robots.txt but the spelling must be exact; a hyphen in the wrong place silently disables the rule. Use the robots.txt generator to pick from this list by category instead of typing tokens by hand.
AI & LLM Bots (88)
Training crawlers, search-and-cite fetchers, on-demand user agents and, new this release, AI coding agents. Browse the full AI & LLM Bots category.
| User-agent token | Operator | Purpose | Safety |
|---|---|---|---|
| AddSearchBot NEW REPORTED | AddSearch | indexing | Safe |
| AI2Bot | Allen Institute for AI | AI training | Safe |
| AI2Bot-DeepResearchEval NEW REPORTED | Allen Institute for AI | on-demand fetch | Safe |
| aiHitBot NEW REPORTED | aiHit | AI training | Caution |
| AIWebIndex NEW | Lyrenth | indexing | Safe |
| amazon-QBusiness NEW REPORTED | Amazon | indexing | Safe |
| AmazonBuyForMe NEW REPORTED | Amazon | on-demand fetch | Safe |
| Amzn-User | Amazon | on-demand fetch | Safe |
| Andibot NEW REPORTED | Andi | training + search | Safe |
| anthropic-ai NEW REPORTED | Anthropic | AI training | Safe |
| atlassian-bot NEW REPORTED | Atlassian | indexing | Safe |
| AzureAI-SearchBot NEW REPORTED | Microsoft | indexing | Safe |
| bedrockbot | Amazon | AI training | Safe |
| Bravebot NEW REPORTED | Brave | on-demand fetch | Safe |
| ChatGLM-Spider NEW REPORTED | Zhipu AI | AI training | Caution |
| ChatGPT Agent NEW | OpenAI | on-demand fetch | Safe |
| ChatGPT Operator NEW | OpenAI | on-demand fetch | Safe |
| ChatGPT-User | OpenAI | on-demand fetch | Safe |
| ChatGPT-User/2.0 | OpenAI | on-demand fetch | Safe |
| Claude-Code NEW REPORTED | Anthropic | on-demand fetch | Safe |
| Claude-SearchBot | Anthropic | on-demand fetch | Safe |
| Claude-User | Anthropic | on-demand fetch | Safe |
| Claude-Web NEW REPORTED | Anthropic | on-demand fetch | Safe |
| ClaudeBot | Anthropic | AI training | Safe |
| Cloudflare-AutoRAG NEW REPORTED | Cloudflare | indexing | Safe |
| cohere-ai NEW REPORTED | Cohere | training + search | Safe |
| cohere-training-data-crawler NEW REPORTED | Cohere | AI training | Caution |
| Cursor NEW REPORTED | Anysphere | on-demand fetch | Safe |
| DeepSeekBot NEW REPORTED | DeepSeek | training + search | Aggressive |
| Devin NEW REPORTED | Cognition AI | on-demand fetch | Safe |
| Diffbot-User NEW | Diffbot | on-demand fetch | Safe |
| DoubaoBot NEW REPORTED | ByteDance | training + search | Caution |
| DuckAssistBot | DuckDuckGo | on-demand fetch | Safe |
| ERNIEBot NEW REPORTED | Baidu | AI training | Caution |
| ExaBot NEW REPORTED | Exa | indexing | Safe |
| ExaSearchBot NEW | Exa | indexing | Safe |
| FirecrawlAgent | Firecrawl | scraping | Caution |
| Google-Extended | AI training | Safe | |
| Google-Firebase NEW REPORTED | on-demand fetch | Safe | |
| Google-Gemini-CLI NEW REPORTED | on-demand fetch | Safe | |
| GoogleAgent-Mariner NEW REPORTED | on-demand fetch | Safe | |
| GoogleAgent-URLContext NEW REPORTED | on-demand fetch | Safe | |
| GPTBot | OpenAI | AI training | Safe |
| GrokBot NEW REPORTED | xAI | training + search | Safe |
| iAskBot NEW REPORTED | iAsk | training + search | Safe |
| iaskspider NEW REPORTED | iAsk | scraping | Safe |
| img2dataset | Open Source Community | AI training | Safe |
| kagi-fetcher NEW REPORTED | Kagi | on-demand fetch | Safe |
| Kimi-SearchBot NEW | Moonshot AI | indexing | Safe |
| Kimi-User | Moonshot AI | on-demand fetch | Safe |
| KimiBot NEW | Moonshot AI | AI training | Caution |
| KlaviyoAIBot NEW | Klaviyo | AI training | Safe |
| LAIONDownloader NEW REPORTED | LAION | AI training | Caution |
| LinerBot NEW | LINER | training + search | Safe |
| LinkupBot NEW | Linkup | indexing | Safe |
| Manus-User NEW REPORTED | Manus | on-demand fetch | Safe |
| Meta-ExternalAgent | Meta | AI training | Caution |
| Meta-ExternalFetcher | Meta | on-demand fetch | Caution |
| meta-webindexer | Meta | indexing | Caution |
| MistralAI-Index | Mistral AI | indexing | Safe |
| MistralAI-Training NEW | Mistral AI | AI training | Safe |
| MistralAI-User | Mistral AI | on-demand fetch | Safe |
| MistralAI-User-1.0 | Mistral AI | on-demand fetch | Safe |
| Mozilla-Tabstack NEW | Mozilla | on-demand fetch | Safe |
| MyCentralAIScraperBot NEW REPORTED | MyCentral AI | AI training | Safe |
| NovaAct NEW REPORTED | Amazon | scraping | Safe |
| OAI-AdsBot | OpenAI | monitoring | Safe |
| OAI-SearchBot | OpenAI | on-demand fetch | Safe |
| opencode NEW REPORTED | OpenCode | on-demand fetch | Safe |
| PanguBot NEW REPORTED | Huawei | AI training | Safe |
| Perplexity-User | Perplexity AI | on-demand fetch | Caution |
| Perplexity-User-1.0 | Perplexity AI | on-demand fetch | Caution |
| PerplexityBot | Perplexity AI | training + search | Safe |
| PhindBot NEW REPORTED | Phind | on-demand fetch | Safe |
| QuillBot NEW REPORTED | QuillBot | on-demand fetch | Caution |
| QwenBot NEW REPORTED | Alibaba | AI training | Caution |
| SBIntuitions-SearchBot NEW | SB Intuitions | on-demand fetch | Safe |
| SBIntuitionsBot NEW | SB Intuitions | AI training | Safe |
| Shap-User NEW REPORTED | Parallel | on-demand fetch | Safe |
| ShapBot NEW | Parallel | indexing | Caution |
| TavilyBot NEW REPORTED | Tavily | on-demand fetch | Safe |
| TerraCotta NEW | Ceramic AI | AI training | Caution |
| TongyiBot NEW REPORTED | Alibaba | on-demand fetch | Caution |
| Trae NEW REPORTED | ByteDance | on-demand fetch | Safe |
| WRTNBot NEW REPORTED | WRTN | training + search | Safe |
| xAI-SearchBot NEW REPORTED | xAI | indexing | Caution |
| YiyanBot NEW REPORTED | Baidu | on-demand fetch | Caution |
| YouBot NEW REPORTED | You | training + search | Safe |
Search Engines (62)
Classic web indexers. Blocking these costs you organic rankings, which is why the checker scores them as critical. Browse the full Search Engines category.
| User-agent token | Operator | Purpose | Safety |
|---|---|---|---|
| AdIdxBot NEW | Microsoft | monitoring | Safe |
| Amzn-SearchBot | Amazon | indexing | Safe |
| Applebot | Apple | indexing | Safe |
| Applebot-Extended | Apple | AI training | Caution |
| baidu | Baidu | indexing | Safe |
| Baiduspider | Baidu | indexing | Safe |
| Baiduspider-ads NEW | Baidu | monitoring | Safe |
| Baiduspider-cpro NEW | Baidu | monitoring | Safe |
| Baiduspider-favo NEW | Baidu | indexing | Safe |
| Baiduspider-image NEW | Baidu | indexing | Safe |
| Baiduspider-news NEW | Baidu | indexing | Safe |
| Baiduspider-video NEW | Baidu | indexing | Safe |
| bingbot | Microsoft | indexing | Safe |
| BingPreview NEW | Microsoft | indexing | Safe |
| BingVideoPreview NEW | Microsoft | indexing | Safe |
| DuckDuckBot | DuckDuckGo | indexing | Safe |
| Googlebot | indexing | Safe | |
| IbouBot NEW REPORTED | Ibou | indexing | Safe |
| Kagibot NEW | Kagi | indexing | Safe |
| MicrosoftPreview NEW | Microsoft | link preview | Safe |
| Mojeek | Mojeek | indexing | Safe |
| MojeekBot | Mojeek | indexing | Safe |
| PetalBot | Huawei | indexing | Safe |
| SeznamBot | Seznam | indexing | Safe |
| Slurp NEW | Yahoo | indexing | Safe |
| YaDirectFetcher NEW | Yandex | monitoring | Safe |
| Yandex | Yandex | indexing | Safe |
| YandexAccessibilityBot NEW | Yandex | indexing | Safe |
| YandexAdditional | Yandex | indexing | Safe |
| YandexAdditionalBot | Yandex | indexing | Safe |
| YandexAdNet NEW | Yandex | monitoring | Safe |
| YandexBlogs NEW | Yandex | indexing | Safe |
| YandexBot | Yandex | indexing | Safe |
| YandexCalendar NEW | Yandex | on-demand fetch | Safe |
| YandexCheckBot NEW | Yandex | monitoring | Safe |
| YandexComBot NEW | Yandex | indexing | Safe |
| YandexDirect NEW | Yandex | monitoring | Safe |
| YandexDirectDyn NEW | Yandex | monitoring | Safe |
| YandexFavicons NEW | Yandex | indexing | Safe |
| YandexFeed NEW | Yandex | indexing | Safe |
| YandexImageResizer NEW | Yandex | on-demand fetch | Safe |
| YandexImages NEW | Yandex | indexing | Safe |
| YandexMarket NEW | Yandex | indexing | Safe |
| YandexMedia NEW | Yandex | indexing | Safe |
| YandexMobileBot NEW | Yandex | indexing | Safe |
| YandexMobileScreenShotBot NEW | Yandex | monitoring | Safe |
| YandexOntoDB NEW | Yandex | indexing | Safe |
| YandexOntoDBAPI NEW | Yandex | on-demand fetch | Safe |
| YandexPagechecker NEW | Yandex | monitoring | Safe |
| YandexPartner NEW | Yandex | monitoring | Safe |
| YandexRCA NEW | Yandex | indexing | Safe |
| YandexRenderResourcesBot NEW | Yandex | indexing | Safe |
| YandexScreenshotBot NEW | Yandex | monitoring | Safe |
| YandexSearchShop NEW | Yandex | indexing | Safe |
| YandexSitelinks NEW | Yandex | monitoring | Safe |
| YandexSpravBot NEW | Yandex | indexing | Safe |
| YandexUserproxy NEW | Yandex | on-demand fetch | Safe |
| YandexVerticals NEW | Yandex | indexing | Safe |
| YandexVertis NEW | Yandex | indexing | Safe |
| YandexVideo NEW | Yandex | indexing | Safe |
| YandexVideoParser NEW | Yandex | indexing | Safe |
| YandexWebmaster NEW | Yandex | monitoring | Safe |
Google Bots (24)
Google's specialised fetchers: Discover, News, Images, Video, the Inspection Tool and the user-triggered fetchers such as Google-GeminiNotebook. Browse the full Google Bots category.
| User-agent token | Operator | Purpose | Safety |
|---|---|---|---|
| AdsBot-Google NEW | monitoring | Safe | |
| AdsBot-Google-Mobile NEW | monitoring | Safe | |
| APIs-Google NEW | on-demand fetch | Safe | |
| FeedFetcher-Google | indexing | Safe | |
| Gemini-Deep-Research NEW REPORTED | on-demand fetch | Safe | |
| Google-Agent | on-demand fetch | Safe | |
| Google-CWS NEW | monitoring | Safe | |
| Google-GeminiNotebook NEW | on-demand fetch | Safe | |
| Google-InspectionTool | monitoring | Safe | |
| Google-NotebookLM NEW REPORTED | on-demand fetch | Safe | |
| Google-Pinpoint NEW | on-demand fetch | Safe | |
| Google-Read-Aloud | on-demand fetch | Safe | |
| Google-Safety NEW | monitoring | Safe | |
| Google-Site-Verification | monitoring | Safe | |
| Googlebot-Image | indexing | Safe | |
| Googlebot-News | indexing | Safe | |
| Googlebot-Video | indexing | Safe | |
| GoogleMessages NEW | link preview | Safe | |
| GoogleOther | indexing | Safe | |
| GoogleOther-Image | indexing | Safe | |
| GoogleOther-Video | indexing | Safe | |
| GoogleProducer | indexing | Safe | |
| Mediapartners-Google NEW | monitoring | Safe | |
| Storebot-Google | indexing | Safe |
SEO Tools (19)
Backlink and rank monitors. Harmless to rankings but heavy on bandwidth on large sites. Browse the full SEO Tools category.
| User-agent token | Operator | Purpose | Safety |
|---|---|---|---|
| AhrefsBot | Ahrefs | monitoring | Safe |
| AhrefsSiteAudit NEW | Ahrefs | monitoring | Safe |
| AwarioBot NEW | Awario | monitoring | Safe |
| AwarioRssBot | Awario | monitoring | Safe |
| AwarioSmartBot | Awario | monitoring | Safe |
| Chrome-Lighthouse NEW REPORTED | monitoring | Safe | |
| DataForSeoBot | DataForSEO | monitoring | Safe |
| DotBot | Moz | monitoring | Caution |
| MJ12bot | Majestic | monitoring | Caution |
| NeticleBot | Neticle | monitoring | Safe |
| Netvibes NEW REPORTED | Dassault Systemes (Netvibes) | monitoring | Safe |
| Rogerbot NEW | Moz | monitoring | Safe |
| Screaming Frog SEO Spider | Screaming Frog | monitoring | Safe |
| SemrushBot | Semrush | monitoring | Safe |
| SemrushBot-OCOB | Semrush | monitoring | Caution |
| SemrushBot-SWA NEW | Semrush | monitoring | Safe |
| Sidetrade Indexer Bot NEW REPORTED | Sidetrade | monitoring | Safe |
| YandexMetrika NEW | Yandex | monitoring | Safe |
| YandexTracker NEW | Yandex | on-demand fetch | Safe |
Social Bots (9)
Link unfurlers that build the preview card when someone shares your URL. Browse the full Social Bots category.
| User-agent token | Operator | Purpose | Safety |
|---|---|---|---|
| FacebookBot NEW REPORTED | Meta | link preview | Safe |
| facebookexternalhit | Meta | link preview | Safe |
| LinkedInBot | link preview | Safe | |
| Meta-ExternalAds NEW | Meta | monitoring | Safe |
| Pinterestbot | link preview | Safe | |
| Quora-Bot NEW REPORTED | Quora | link preview | Safe |
| Slackbot | Slack | link preview | Safe |
| Slackbot-LinkExpanding NEW | Slack | link preview | Safe |
| Twitterbot | X (Twitter) | link preview | Safe |
Scrapers (39)
Dataset builders, archive crawlers and commercial extraction services. The category with the most aggressive tokens. Browse the full Scrapers category.
| User-agent token | Operator | Purpose | Safety |
|---|---|---|---|
| Amazon Kendra | Amazon | scraping | Safe |
| ApifyBot NEW REPORTED | Apify | scraping | Caution |
| Barkrowler | Babbar | scraping | Safe |
| BLEXBot NEW REPORTED | WebMeUp | scraping | Safe |
| Brightbot NEW | Bright Data | AI training | Aggressive |
| Bytespider NEW | ByteDance | training + search | Aggressive |
| CCBot | Common Crawl | AI training | Caution |
| coccocbot-web | Unknown | scraping | Safe |
| Cotoyogi NEW | ROIS-DS | AI training | Safe |
| Crawl4AI | Unknown | scraping | Safe |
| crawler4j | Unknown | scraping | Safe |
| Crawlspace NEW REPORTED | Crawlspace | scraping | Safe |
| Diffbot | Diffbot | scraping | Safe |
| Echobot Bot NEW REPORTED | Dealfront | scraping | Safe |
| EchoboxBot NEW REPORTED | Echobox | scraping | Safe |
| Factset_spyderbot NEW REPORTED | FactSet | scraping | Caution |
| FriendlyCrawler NEW REPORTED | FriendlyCrawler | scraping | Safe |
| ICC-Crawler | ICC | scraping | Caution |
| ImagesiftBot | Hive | scraping | Safe |
| imgproxy | Open Source | scraping | Safe |
| ISSCyberRiskCrawler NEW REPORTED | ISS | scraping | Safe |
| Kangaroo Bot NEW REPORTED | Kangaroo | scraping | Safe |
| magpie-crawler | Brandwatch | scraping | Caution |
| MauiBot NEW REPORTED | MauiBot | scraping | Safe |
| netEstate Imprint Crawler NEW REPORTED | netEstate | scraping | Safe |
| news-please | Open Source | scraping | Aggressive |
| omgili | Unknown | scraping | Safe |
| omgilibot | Unknown | scraping | Safe |
| Panscient NEW | Panscient | scraping | Safe |
| Poseidon Research Crawler NEW REPORTED | Poseidon Research | scraping | Safe |
| QualifiedBot NEW REPORTED | Qualified | scraping | Safe |
| Scrapy | Unknown | scraping | Safe |
| SeekportBot | Seekport | scraping | Safe |
| TikTokSpider NEW REPORTED | ByteDance | AI training | Aggressive |
| Timpibot NEW REPORTED | Timpi | scraping | Safe |
| VelenPublicWebCrawler | Velen | scraping | Safe |
| Webzio-Extended NEW | Webz.io | AI training | Caution |
| yacy | YaCy Community | indexing | Safe |
| yacybot | YaCy Community | indexing | Safe |
Cloud Services (3)
Agents run by cloud platforms on behalf of their customers. Browse the full Cloud Services category.
| User-agent token | Operator | Purpose | Safety |
|---|---|---|---|
| AliyunSecBot | Alibaba Cloud | monitoring | Safe |
| Amazonbot | Amazon | AI training | Safe |
| Google-CloudVertexBot | on-demand fetch | Safe |
Other Agents (4)
Archive and academic-integrity crawlers that fit nowhere else. Browse the full Other Agents category.
| User-agent token | Operator | Purpose | Safety |
|---|---|---|---|
| archive.org_bot | Internet Archive | scraping | Caution |
| Arquivo-web-crawler NEW | Arquivo.pt | scraping | Safe |
| ia_archiver NEW REPORTED | Internet Archive | scraping | Caution |
| TurnitinBot | Turnitin | scraping | Safe |
Turning the Roster into a Policy
The hybrid approach most publishers land on is simple: allow everything in Search Engines, Google Bots and Social Bots without exception; allow the on-demand and search-and-cite AI fetchers so you appear in answers; decide deliberately about pure training crawlers; and rate-limit or block the Aggressive scrapers at the firewall rather than in robots.txt, because they do not read it. Our guide to blocking training while allowing search gives the exact file.
# Training crawlers: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
User-agent: Applebot-Extended
User-agent: Brightbot
Disallow: /
# Search-and-cite and on-demand fetchers: allow
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Google-GeminiNotebook
Allow: /
After editing, validate the file with the robots.txt validator, which parses it against all 248 tokens and flags any retired ones, then re-run the free AI crawler check to confirm the live result. Remember that robots.txt is advisory. The Invisible AI Crawler Block case study shows a site whose robots.txt was perfect and that still blocked every AI crawler at the edge.
Run the Free Check
Run a free AI crawler check on your website to see which of the 88 AI crawlers can access your content. The tool analyzes your robots.txt, looks for an llms.txt file, checks for firewall blocks, and gives you an AI Visibility Score from 0 to 100. You can also explore the full AI bot directory or run a deeper GEO Audit.
How to Verify a Bot Is Who It Says It Is
Bad actors spoof popular tokens to disguise scraping, and Googlebot and GPTBot are the two most impersonated. Genuine major crawlers publish IP ranges or support reverse DNS: capture the source IP from your logs, resolve it, confirm the hostname belongs to the operator, then forward-confirm the IP. Cross-check against the published range where one exists, and block confirmed impostors at the firewall so the real crawler is unaffected. See verify AI bots with reverse DNS and spotting spoofed user-agents.
Where to Go From Here
This roster is one piece of a larger picture. The full directory documents every crawler with its operator, purpose and safety rating; the operators index groups them by company so you can see every token a single vendor runs; and the batch URL checker audits many sites in one pass if you manage a portfolio. For what is coming next, read the 2026 watchlist.
Your Action Checklist
- □Establish a baseline with a free AI crawler check and write down the score before you change anything.
- □Search your robots.txt for the retired tokens above and replace each with its successor.
- □Decide one policy per category, not one per token, and write it down.
- □Move Aggressive scrapers out of robots.txt and into a firewall rule.
- □Re-check after every CDN, WAF or platform change; the edge can block what robots.txt allows.