This is the full roster of every crawler, fetcher and agent that AI Crawler Check tests your site against: 248 bots across 8 categories, operated by 103 companies and projects. Each token links to its own directory profile with the exact robots.txt rule, the operator's documentation and a safety verdict.

The list is generated from the same database the checker runs on, so it cannot drift from what the tool actually detects. It was last rebuilt for release 6.1, which checked every token against its operator's own documentation, labelled each one Official or Reported, and removed the strings that turned out to be dead spellings or bare product names. Before reading on, run a free AI crawler check so you know which of these can currently reach you.

Key Takeaways

  • 88 of the 248 are AI and LLM crawlers. The rest are search engines, Google fetchers, SEO tools, social unfurlers, scrapers and cloud agents, and they need different rules.
  • 211 tokens are rated Safe, 32 Caution and 5 Aggressive. Aggressive means documented evidence of ignoring robots.txt or excessive crawl rates.
  • 52 tokens are user-triggered fetchers (ChatGPT-User, Claude-User, Google-GeminiNotebook, Perplexity-User, Kimi-User). They fetch a page because a person asked for it and, by their operators' own statements, may not honour robots.txt.
  • 177 are named on their operator's own documentation and 71 are Reported: named by a credible public registry but not by the operator. Every profile shows which, and links the source, so you can decide how much weight to give the rule.
How AI crawlers work: pipeline from website to AI answer Flow diagram showing four stages: your website is checked against robots.txt rules, an AI crawler reads allowed content, the content feeds an AI model, and the model produces AI answers that can cite your site. Your Website pages + content example.com robots.txt User-agent: GPTBot Allow: / the gatekeeper AI Crawler GPTBot, ClaudeBot, PerplexityBot... reads allowed pages AI Model training + retrieval AI Answers citing your site Block the crawler at step 2 and your content never reaches the AI answer at step 4.
How AI crawlers work: from your website through robots.txt to AI-generated answers.

The Shape of the Database

Not every bot is an AI bot, and treating them all the same is the most common mistake we see in robots.txt files. A rule that blocks Googlebot because it sat in the same block as GPTBot costs you organic rankings for nothing. The eight categories below exist so that you can write one policy per class of traffic rather than one per token.

Bots per category
AI & LLM Bots 88
Search Engines 62
Google Bots 24
SEO Tools 19
Social Bots 9
Scrapers 39
Cloud Services 3
Other Agents 4

Counts are generated from the live database at build time and sum to 248.

By purpose, 28 tokens exist to collect training data, 52 fetch a page on a user's request (these usually ignore robots.txt by design, because a human asked for the page), 67 build a search or answer index, 11 do both training and search, and the remainder scrape, monitor or unfurl links. The distinction matters because the on-demand fetchers are the ones that put your page in front of a person in real time; blocking them removes you from the answer while the training crawlers keep reading. Our guide to training bots vs search bots walks through the policy implications.

New in Release 6.1: 155 Additions

Additions come in two tiers, and every profile wears its tier as a badge. Official docs (177 of them) means the operator's own crawler page names the token and the profile links to it. Reported (71 of them) means no operator page exists but a credible public source does (the ai.robots.txt registry, Trakkr's observatory, user-agents.net); the profile says so plainly and links that source. Nothing is listed on hearsay alone, and bare product names such as Grok, OpenAI or Operator are not tokens and are not listed.

Google documented fetchers

Moonshot AI (Kimi)

New fetchers from known operators

Restored after the second verification pass (each is on its vendor page)

Yandex robots, from the official Yandex list

Baidu product spiders

Google special-case crawlers

Bing crawlers

Search, AI and SEO operators with documented sibling tokens

ChatGPT Agent, confirmed by OpenAI's allowlisting guide (a signed agent, not a robots.txt token)

Vendors found through the Trakkr cross-check whose own pages print the token

Reported tokens: no operator page, but a credible public registry names them

Official, Reported, Retired: How to Read the Badges

Most crawler lists mix three very different things: tokens the vendor publishes, tokens that only registries and server logs know, and dead spellings that match nothing. We separate them, because the right action is different for each.

Official docs (177)

The operator's own page names the token. Purpose, robots.txt behaviour and, where published, IP ranges come from that page. OpenAI, Anthropic, Google, Meta, Apple, Amazon, Microsoft, Mistral, Perplexity, Moonshot, Yandex, Baidu and ByteDance's Toutiao Search team all publish such pages.

Reported (71)

No operator page, but a credible public source names it. Claude-Web, anthropic-ai, GrokBot, xAI-SearchBot, DeepSeekBot, QwenBot, cohere-ai, Gemini-Deep-Research and the AI coding agents sit here. The rule on the profile is the best available spelling; treat it as best effort and pair it with a firewall rule if the traffic matters.

Retired or never a token

  • anthropic-ai replaced by ClaudeBot
  • claude-web replaced by ClaudeBot
  • google-notebooklm replaced by Google-GeminiNotebook
  • notebooklm replaced by Google-GeminiNotebook
  • chatgpt operator replaced by ChatGPT-User
  • chatgpt agent replaced by ChatGPT-User
  • operator replaced by ChatGPT-User
  • teoma

The validator flags these and names the token to add. ChatGPT Agent and ChatGPT Operator are real, but they are CDN bot-manager names for OpenAI's signed Cloud browser, not robots.txt tokens; their profiles explain Web Bot Auth.

Deliberately absent: dead Yahoo spellings (YahooSeeker, Yahoo-MMCrawler, Yahoo-Blogs, Yahoo-FeedSeeker), Teoma, Ai2Bot-Dolma (AI2's page names only AI2Bot), and bare company or product names that were never user-agent tokens. The full exclusion list, with a reason for each, ships with the site's source data.

The Full Roster, by Category

Tokens are the exact strings to use in User-agent: lines. Matching is case-insensitive in robots.txt but the spelling must be exact; a hyphen in the wrong place silently disables the rule. Use the robots.txt generator to pick from this list by category instead of typing tokens by hand.

AI & LLM Bots (88)

Training crawlers, search-and-cite fetchers, on-demand user agents and, new this release, AI coding agents. Browse the full AI & LLM Bots category.

User-agent tokenOperatorPurposeSafety
AddSearchBot NEW REPORTEDAddSearchindexingSafe
AI2BotAllen Institute for AIAI trainingSafe
AI2Bot-DeepResearchEval NEW REPORTEDAllen Institute for AIon-demand fetchSafe
aiHitBot NEW REPORTEDaiHitAI trainingCaution
AIWebIndex NEWLyrenthindexingSafe
amazon-QBusiness NEW REPORTEDAmazonindexingSafe
AmazonBuyForMe NEW REPORTEDAmazonon-demand fetchSafe
Amzn-UserAmazonon-demand fetchSafe
Andibot NEW REPORTEDAnditraining + searchSafe
anthropic-ai NEW REPORTEDAnthropicAI trainingSafe
atlassian-bot NEW REPORTEDAtlassianindexingSafe
AzureAI-SearchBot NEW REPORTEDMicrosoftindexingSafe
bedrockbotAmazonAI trainingSafe
Bravebot NEW REPORTEDBraveon-demand fetchSafe
ChatGLM-Spider NEW REPORTEDZhipu AIAI trainingCaution
ChatGPT Agent NEWOpenAIon-demand fetchSafe
ChatGPT Operator NEWOpenAIon-demand fetchSafe
ChatGPT-UserOpenAIon-demand fetchSafe
ChatGPT-User/2.0OpenAIon-demand fetchSafe
Claude-Code NEW REPORTEDAnthropicon-demand fetchSafe
Claude-SearchBotAnthropicon-demand fetchSafe
Claude-UserAnthropicon-demand fetchSafe
Claude-Web NEW REPORTEDAnthropicon-demand fetchSafe
ClaudeBotAnthropicAI trainingSafe
Cloudflare-AutoRAG NEW REPORTEDCloudflareindexingSafe
cohere-ai NEW REPORTEDCoheretraining + searchSafe
cohere-training-data-crawler NEW REPORTEDCohereAI trainingCaution
Cursor NEW REPORTEDAnysphereon-demand fetchSafe
DeepSeekBot NEW REPORTEDDeepSeektraining + searchAggressive
Devin NEW REPORTEDCognition AIon-demand fetchSafe
Diffbot-User NEWDiffboton-demand fetchSafe
DoubaoBot NEW REPORTEDByteDancetraining + searchCaution
DuckAssistBotDuckDuckGoon-demand fetchSafe
ERNIEBot NEW REPORTEDBaiduAI trainingCaution
ExaBot NEW REPORTEDExaindexingSafe
ExaSearchBot NEWExaindexingSafe
FirecrawlAgentFirecrawlscrapingCaution
Google-ExtendedGoogleAI trainingSafe
Google-Firebase NEW REPORTEDGoogleon-demand fetchSafe
Google-Gemini-CLI NEW REPORTEDGoogleon-demand fetchSafe
GoogleAgent-Mariner NEW REPORTEDGoogleon-demand fetchSafe
GoogleAgent-URLContext NEW REPORTEDGoogleon-demand fetchSafe
GPTBotOpenAIAI trainingSafe
GrokBot NEW REPORTEDxAItraining + searchSafe
iAskBot NEW REPORTEDiAsktraining + searchSafe
iaskspider NEW REPORTEDiAskscrapingSafe
img2datasetOpen Source CommunityAI trainingSafe
kagi-fetcher NEW REPORTEDKagion-demand fetchSafe
Kimi-SearchBot NEWMoonshot AIindexingSafe
Kimi-UserMoonshot AIon-demand fetchSafe
KimiBot NEWMoonshot AIAI trainingCaution
KlaviyoAIBot NEWKlaviyoAI trainingSafe
LAIONDownloader NEW REPORTEDLAIONAI trainingCaution
LinerBot NEWLINERtraining + searchSafe
LinkupBot NEWLinkupindexingSafe
Manus-User NEW REPORTEDManuson-demand fetchSafe
Meta-ExternalAgentMetaAI trainingCaution
Meta-ExternalFetcherMetaon-demand fetchCaution
meta-webindexerMetaindexingCaution
MistralAI-IndexMistral AIindexingSafe
MistralAI-Training NEWMistral AIAI trainingSafe
MistralAI-UserMistral AIon-demand fetchSafe
MistralAI-User-1.0Mistral AIon-demand fetchSafe
Mozilla-Tabstack NEWMozillaon-demand fetchSafe
MyCentralAIScraperBot NEW REPORTEDMyCentral AIAI trainingSafe
NovaAct NEW REPORTEDAmazonscrapingSafe
OAI-AdsBotOpenAImonitoringSafe
OAI-SearchBotOpenAIon-demand fetchSafe
opencode NEW REPORTEDOpenCodeon-demand fetchSafe
PanguBot NEW REPORTEDHuaweiAI trainingSafe
Perplexity-UserPerplexity AIon-demand fetchCaution
Perplexity-User-1.0Perplexity AIon-demand fetchCaution
PerplexityBotPerplexity AItraining + searchSafe
PhindBot NEW REPORTEDPhindon-demand fetchSafe
QuillBot NEW REPORTEDQuillBoton-demand fetchCaution
QwenBot NEW REPORTEDAlibabaAI trainingCaution
SBIntuitions-SearchBot NEWSB Intuitionson-demand fetchSafe
SBIntuitionsBot NEWSB IntuitionsAI trainingSafe
Shap-User NEW REPORTEDParallelon-demand fetchSafe
ShapBot NEWParallelindexingCaution
TavilyBot NEW REPORTEDTavilyon-demand fetchSafe
TerraCotta NEWCeramic AIAI trainingCaution
TongyiBot NEW REPORTEDAlibabaon-demand fetchCaution
Trae NEW REPORTEDByteDanceon-demand fetchSafe
WRTNBot NEW REPORTEDWRTNtraining + searchSafe
xAI-SearchBot NEW REPORTEDxAIindexingCaution
YiyanBot NEW REPORTEDBaiduon-demand fetchCaution
YouBot NEW REPORTEDYoutraining + searchSafe

Search Engines (62)

Classic web indexers. Blocking these costs you organic rankings, which is why the checker scores them as critical. Browse the full Search Engines category.

User-agent tokenOperatorPurposeSafety
AdIdxBot NEWMicrosoftmonitoringSafe
Amzn-SearchBotAmazonindexingSafe
ApplebotAppleindexingSafe
Applebot-ExtendedAppleAI trainingCaution
baiduBaiduindexingSafe
BaiduspiderBaiduindexingSafe
Baiduspider-ads NEWBaidumonitoringSafe
Baiduspider-cpro NEWBaidumonitoringSafe
Baiduspider-favo NEWBaiduindexingSafe
Baiduspider-image NEWBaiduindexingSafe
Baiduspider-news NEWBaiduindexingSafe
Baiduspider-video NEWBaiduindexingSafe
bingbotMicrosoftindexingSafe
BingPreview NEWMicrosoftindexingSafe
BingVideoPreview NEWMicrosoftindexingSafe
DuckDuckBotDuckDuckGoindexingSafe
GooglebotGoogleindexingSafe
IbouBot NEW REPORTEDIbouindexingSafe
Kagibot NEWKagiindexingSafe
MicrosoftPreview NEWMicrosoftlink previewSafe
MojeekMojeekindexingSafe
MojeekBotMojeekindexingSafe
PetalBotHuaweiindexingSafe
SeznamBotSeznamindexingSafe
Slurp NEWYahooindexingSafe
YaDirectFetcher NEWYandexmonitoringSafe
YandexYandexindexingSafe
YandexAccessibilityBot NEWYandexindexingSafe
YandexAdditionalYandexindexingSafe
YandexAdditionalBotYandexindexingSafe
YandexAdNet NEWYandexmonitoringSafe
YandexBlogs NEWYandexindexingSafe
YandexBotYandexindexingSafe
YandexCalendar NEWYandexon-demand fetchSafe
YandexCheckBot NEWYandexmonitoringSafe
YandexComBot NEWYandexindexingSafe
YandexDirect NEWYandexmonitoringSafe
YandexDirectDyn NEWYandexmonitoringSafe
YandexFavicons NEWYandexindexingSafe
YandexFeed NEWYandexindexingSafe
YandexImageResizer NEWYandexon-demand fetchSafe
YandexImages NEWYandexindexingSafe
YandexMarket NEWYandexindexingSafe
YandexMedia NEWYandexindexingSafe
YandexMobileBot NEWYandexindexingSafe
YandexMobileScreenShotBot NEWYandexmonitoringSafe
YandexOntoDB NEWYandexindexingSafe
YandexOntoDBAPI NEWYandexon-demand fetchSafe
YandexPagechecker NEWYandexmonitoringSafe
YandexPartner NEWYandexmonitoringSafe
YandexRCA NEWYandexindexingSafe
YandexRenderResourcesBot NEWYandexindexingSafe
YandexScreenshotBot NEWYandexmonitoringSafe
YandexSearchShop NEWYandexindexingSafe
YandexSitelinks NEWYandexmonitoringSafe
YandexSpravBot NEWYandexindexingSafe
YandexUserproxy NEWYandexon-demand fetchSafe
YandexVerticals NEWYandexindexingSafe
YandexVertis NEWYandexindexingSafe
YandexVideo NEWYandexindexingSafe
YandexVideoParser NEWYandexindexingSafe
YandexWebmaster NEWYandexmonitoringSafe

Google Bots (24)

Google's specialised fetchers: Discover, News, Images, Video, the Inspection Tool and the user-triggered fetchers such as Google-GeminiNotebook. Browse the full Google Bots category.

User-agent tokenOperatorPurposeSafety
AdsBot-Google NEWGooglemonitoringSafe
AdsBot-Google-Mobile NEWGooglemonitoringSafe
APIs-Google NEWGoogleon-demand fetchSafe
FeedFetcher-GoogleGoogleindexingSafe
Gemini-Deep-Research NEW REPORTEDGoogleon-demand fetchSafe
Google-AgentGoogleon-demand fetchSafe
Google-CWS NEWGooglemonitoringSafe
Google-GeminiNotebook NEWGoogleon-demand fetchSafe
Google-InspectionToolGooglemonitoringSafe
Google-NotebookLM NEW REPORTEDGoogleon-demand fetchSafe
Google-Pinpoint NEWGoogleon-demand fetchSafe
Google-Read-AloudGoogleon-demand fetchSafe
Google-Safety NEWGooglemonitoringSafe
Google-Site-VerificationGooglemonitoringSafe
Googlebot-ImageGoogleindexingSafe
Googlebot-NewsGoogleindexingSafe
Googlebot-VideoGoogleindexingSafe
GoogleMessages NEWGooglelink previewSafe
GoogleOtherGoogleindexingSafe
GoogleOther-ImageGoogleindexingSafe
GoogleOther-VideoGoogleindexingSafe
GoogleProducerGoogleindexingSafe
Mediapartners-Google NEWGooglemonitoringSafe
Storebot-GoogleGoogleindexingSafe

SEO Tools (19)

Backlink and rank monitors. Harmless to rankings but heavy on bandwidth on large sites. Browse the full SEO Tools category.

User-agent tokenOperatorPurposeSafety
AhrefsBotAhrefsmonitoringSafe
AhrefsSiteAudit NEWAhrefsmonitoringSafe
AwarioBot NEWAwariomonitoringSafe
AwarioRssBotAwariomonitoringSafe
AwarioSmartBotAwariomonitoringSafe
Chrome-Lighthouse NEW REPORTEDGooglemonitoringSafe
DataForSeoBotDataForSEOmonitoringSafe
DotBotMozmonitoringCaution
MJ12botMajesticmonitoringCaution
NeticleBotNeticlemonitoringSafe
Netvibes NEW REPORTEDDassault Systemes (Netvibes)monitoringSafe
Rogerbot NEWMozmonitoringSafe
Screaming Frog SEO SpiderScreaming FrogmonitoringSafe
SemrushBotSemrushmonitoringSafe
SemrushBot-OCOBSemrushmonitoringCaution
SemrushBot-SWA NEWSemrushmonitoringSafe
Sidetrade Indexer Bot NEW REPORTEDSidetrademonitoringSafe
YandexMetrika NEWYandexmonitoringSafe
YandexTracker NEWYandexon-demand fetchSafe

Social Bots (9)

Link unfurlers that build the preview card when someone shares your URL. Browse the full Social Bots category.

User-agent tokenOperatorPurposeSafety
FacebookBot NEW REPORTEDMetalink previewSafe
facebookexternalhitMetalink previewSafe
LinkedInBotLinkedInlink previewSafe
Meta-ExternalAds NEWMetamonitoringSafe
PinterestbotPinterestlink previewSafe
Quora-Bot NEW REPORTEDQuoralink previewSafe
SlackbotSlacklink previewSafe
Slackbot-LinkExpanding NEWSlacklink previewSafe
TwitterbotX (Twitter)link previewSafe

Scrapers (39)

Dataset builders, archive crawlers and commercial extraction services. The category with the most aggressive tokens. Browse the full Scrapers category.

User-agent tokenOperatorPurposeSafety
Amazon KendraAmazonscrapingSafe
ApifyBot NEW REPORTEDApifyscrapingCaution
BarkrowlerBabbarscrapingSafe
BLEXBot NEW REPORTEDWebMeUpscrapingSafe
Brightbot NEWBright DataAI trainingAggressive
Bytespider NEWByteDancetraining + searchAggressive
CCBotCommon CrawlAI trainingCaution
coccocbot-webUnknownscrapingSafe
Cotoyogi NEWROIS-DSAI trainingSafe
Crawl4AIUnknownscrapingSafe
crawler4jUnknownscrapingSafe
Crawlspace NEW REPORTEDCrawlspacescrapingSafe
DiffbotDiffbotscrapingSafe
Echobot Bot NEW REPORTEDDealfrontscrapingSafe
EchoboxBot NEW REPORTEDEchoboxscrapingSafe
Factset_spyderbot NEW REPORTEDFactSetscrapingCaution
FriendlyCrawler NEW REPORTEDFriendlyCrawlerscrapingSafe
ICC-CrawlerICCscrapingCaution
ImagesiftBotHivescrapingSafe
imgproxyOpen SourcescrapingSafe
ISSCyberRiskCrawler NEW REPORTEDISSscrapingSafe
Kangaroo Bot NEW REPORTEDKangarooscrapingSafe
magpie-crawlerBrandwatchscrapingCaution
MauiBot NEW REPORTEDMauiBotscrapingSafe
netEstate Imprint Crawler NEW REPORTEDnetEstatescrapingSafe
news-pleaseOpen SourcescrapingAggressive
omgiliUnknownscrapingSafe
omgilibotUnknownscrapingSafe
Panscient NEWPanscientscrapingSafe
Poseidon Research Crawler NEW REPORTEDPoseidon ResearchscrapingSafe
QualifiedBot NEW REPORTEDQualifiedscrapingSafe
ScrapyUnknownscrapingSafe
SeekportBotSeekportscrapingSafe
TikTokSpider NEW REPORTEDByteDanceAI trainingAggressive
Timpibot NEW REPORTEDTimpiscrapingSafe
VelenPublicWebCrawlerVelenscrapingSafe
Webzio-Extended NEWWebz.ioAI trainingCaution
yacyYaCy CommunityindexingSafe
yacybotYaCy CommunityindexingSafe

Cloud Services (3)

Agents run by cloud platforms on behalf of their customers. Browse the full Cloud Services category.

User-agent tokenOperatorPurposeSafety
AliyunSecBotAlibaba CloudmonitoringSafe
AmazonbotAmazonAI trainingSafe
Google-CloudVertexBotGoogleon-demand fetchSafe

Other Agents (4)

Archive and academic-integrity crawlers that fit nowhere else. Browse the full Other Agents category.

User-agent tokenOperatorPurposeSafety
archive.org_botInternet ArchivescrapingCaution
Arquivo-web-crawler NEWArquivo.ptscrapingSafe
ia_archiver NEW REPORTEDInternet ArchivescrapingCaution
TurnitinBotTurnitinscrapingSafe

Turning the Roster into a Policy

The hybrid approach most publishers land on is simple: allow everything in Search Engines, Google Bots and Social Bots without exception; allow the on-demand and search-and-cite AI fetchers so you appear in answers; decide deliberately about pure training crawlers; and rate-limit or block the Aggressive scrapers at the firewall rather than in robots.txt, because they do not read it. Our guide to blocking training while allowing search gives the exact file.

robots.txt (block training, allow search-and-cite)
# Training crawlers: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
User-agent: Applebot-Extended
User-agent: Brightbot
Disallow: /

# Search-and-cite and on-demand fetchers: allow
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Google-GeminiNotebook
Allow: /

After editing, validate the file with the robots.txt validator, which parses it against all 248 tokens and flags any retired ones, then re-run the free AI crawler check to confirm the live result. Remember that robots.txt is advisory. The Invisible AI Crawler Block case study shows a site whose robots.txt was perfect and that still blocked every AI crawler at the edge.

Run the Free Check

Run a free AI crawler check on your website to see which of the 88 AI crawlers can access your content. The tool analyzes your robots.txt, looks for an llms.txt file, checks for firewall blocks, and gives you an AI Visibility Score from 0 to 100. You can also explore the full AI bot directory or run a deeper GEO Audit.

How to Verify a Bot Is Who It Says It Is

Bad actors spoof popular tokens to disguise scraping, and Googlebot and GPTBot are the two most impersonated. Genuine major crawlers publish IP ranges or support reverse DNS: capture the source IP from your logs, resolve it, confirm the hostname belongs to the operator, then forward-confirm the IP. Cross-check against the published range where one exists, and block confirmed impostors at the firewall so the real crawler is unaffected. See verify AI bots with reverse DNS and spotting spoofed user-agents.

Where to Go From Here

This roster is one piece of a larger picture. The full directory documents every crawler with its operator, purpose and safety rating; the operators index groups them by company so you can see every token a single vendor runs; and the batch URL checker audits many sites in one pass if you manage a portfolio. For what is coming next, read the 2026 watchlist.

Your Action Checklist

  • Establish a baseline with a free AI crawler check and write down the score before you change anything.
  • Search your robots.txt for the retired tokens above and replace each with its successor.
  • Decide one policy per category, not one per token, and write it down.
  • Move Aggressive scrapers out of robots.txt and into a firewall rule.
  • Re-check after every CDN, WAF or platform change; the edge can block what robots.txt allows.