How many of the web's biggest sites block AI crawlers in 2026, and which ones? We fetched the robots.txt file of every domain in the Tranco top 10,000 on 2026-10-05 and parsed 4,962 valid files the way RFC 9309 says a crawler must. This page shows what we found, how we did it, and where the method has limits. The raw counts are free to download as a CSV.
What changed in this update
- : First published. Data fetched the same day.
Key Takeaways
- 14.3% of sites block GPTBot, the most named AI crawler. CCBot (15.4%) and Bytespider (15%) are blocked slightly more often.
- 19.2% block at least one AI training crawler, but only 10.8% block an AI search crawler.
- About half of GPTBot blockers keep OpenAI search open. 347 of 710 still allow OAI-SearchBot.
- 79.9% of Google-Extended blockers keep Googlebot, which is exactly how Google designed the token.
- The top 1,000 sites block AI crawlers about 1.5 times as often as sites ranked 5,001 to 10,000.
- 398 sites answered with an HTML page instead of robots.txt rules, usually a challenge or error page.
Key findings at a glance
Here are the headline numbers. Every percentage uses the same denominator, the 4,962 domains that returned a real robots.txt file, unless the row says otherwise. We explain the denominator in detail in the limitations section, because it changes how you should read the rest of the page.
| Measure | Count | Share |
|---|---|---|
| Domains in the list | 10,000 | 100% |
| Domains that returned any HTTP response | 7,315 | 73.2% |
| Valid robots.txt files parsed | 4,962 | 49.6% of the list |
| Block at least one AI training crawler | 952 | 19.2% |
| Block at least one AI search crawler | 534 | 10.8% |
Block every crawler with User-agent: * and Disallow: / | 177 | 3.6% |
| List a sitemap | 3,341 | 67.3% |
| Use Crawl-delay anywhere | 573 | 11.5% |
| Use a Content-Signal line | 125 | 2.5% |
"Training crawler" here means at least one of GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Meta-ExternalAgent or Applebot-Extended. "AI search crawler" means at least one of OAI-SearchBot, Claude-SearchBot or PerplexityBot. We chose these groups because each operator's own documentation describes the token that way, as we explain below.
Which AI crawlers are blocked most often?
The chart counts a token as blocked when the rules that apply to it stop it from fetching the homepage. That can happen in two ways: the file names the token in its own group, or the file has no group for it and the wildcard group blocks everyone. We show both, because the difference matters when you write your own file.
| Token | Named in the file | Blocked (any group) | Blocked by its own group | Blocked only by * |
|---|---|---|---|---|
| CCBot | 718 (14.5%) | 766 (15.4%) | 588 | 178 |
| Bytespider | 638 (12.9%) | 746 (15%) | 565 | 181 |
| GPTBot | 839 (16.9%) | 710 (14.3%) | 549 | 161 |
| ClaudeBot | 710 (14.3%) | 665 (13.4%) | 494 | 171 |
| Google-Extended | 655 (13.2%) | 597 (12%) | 431 | 166 |
| Meta-ExternalAgent | 506 (10.2%) | 593 (12%) | 415 | 178 |
| Applebot-Extended | 474 (9.6%) | 557 (11.2%) | 383 | 174 |
| Amazonbot | 475 (9.6%) | 545 (11%) | 370 | 175 |
| PerplexityBot | 613 (12.4%) | 520 (10.5%) | 356 | 164 |
| ChatGPT-User | 566 (11.4%) | 450 (9.1%) | 287 | 163 |
| OAI-SearchBot | 495 (10%) | 367 (7.4%) | 205 | 162 |
| Claude-SearchBot | 303 (6.1%) | 361 (7.3%) | 186 | 175 |
| GrokBot | 47 (0.9%) | 236 (4.8%) | 43 | 193 |
Three things stand out. First, the old training crawlers sit at the top: Common Crawl's CCBot and ByteDance's Bytespider were the first bots many site owners heard about, and their blocks have stayed. Second, GPTBot is the token people write most often (16.9% of files name it), yet it is not the most blocked, because some files name it only to allow it. Third, the search crawlers in teal are blocked about half as often as the training crawlers from the same companies.
Why GrokBot looks high
Only 47 files name GrokBot, but 236 block it. The difference, 193 files, is sites that block every crawler with the wildcard group. xAI has not documented a crawler, so few owners write a rule for it. See our GrokBot profile for what is and is not known.
Training bots vs search bots: the split most sites now make
OpenAI, Anthropic and Google each publish separate controls for training and for search. OpenAI's crawler page says each setting "is independent of the others", so a site can allow OAI-SearchBot to appear in ChatGPT search while disallowing GPTBot for training. Anthropic's help page describes ClaudeBot for training and Claude-SearchBot for search. Google's crawler list says Google-Extended "does not impact a site's inclusion in Google Search".
So we asked a simple question: when a site blocks the training token, does it keep the search token open? The answer is "about half the time" for OpenAI and Anthropic, and "almost always" for Google.
| Company | Sites blocking the training token | Of those, search token still allowed | Share that split the two |
|---|---|---|---|
| OpenAI (GPTBot vs OAI-SearchBot) | 710 | 347 | 48.9% |
| Anthropic (ClaudeBot vs Claude-SearchBot) | 665 | 307 | 46.2% |
| Google (Google-Extended vs Googlebot) | 597 | 477 | 79.9% |
The Google number is high because almost nobody blocks Googlebot on purpose: only 11 files block it in its own group. The OpenAI and Anthropic numbers are the interesting ones. Roughly half of the sites that opted out of training also opted out of AI search, often with one copy-pasted block list. If your goal was "no training, but please cite me", a block list like that works against you. Our robots.txt best practices guide shows the split template.
Bigger sites block AI crawlers more often
We split the list into three bands by Tranco rank. Each band has its own denominator: 505 valid files in the top 1,000, 1,625 in ranks 1,001 to 5,000, and 2,832 in ranks 5,001 to 10,000.
| Token | Top 1,000 | Rank 1,001 to 5,000 | Rank 5,001 to 10,000 |
|---|---|---|---|
| GPTBot | 20.8% | 14% | 13.3% |
| ClaudeBot | 20.2% | 13.3% | 12.3% |
| Google-Extended | 18.8% | 11.8% | 11% |
| CCBot | 23% | 14.8% | 14.5% |
| Bytespider | 23.8% | 14.3% | 13.9% |
| OAI-SearchBot | 12.1% | 7.1% | 6.7% |
| PerplexityBot | 17.4% | 10.2% | 9.4% |
| Googlebot | 3.4% | 2.6% | 2.3% |
For every AI token, the top band blocks the most. GPTBot goes from 13.3% in the lowest band to 20.8% in the top 1,000. Cloudflare saw the same pattern in 2024: "the higher-ranked (more popular) an Internet property is, the more likely it is to be targeted by AI bots, and correspondingly, the more likely it is to block such requests" (Cloudflare, July 2024). Big publishers have content licensing deals to protect and legal teams who read the terms. Smaller sites more often want the visibility.
Named is not the same as blocked
A common mistake when reading robots.txt surveys is to count every mention of a bot as a block. Many files name a crawler only to allow it, or to give it a crawl delay. And many files block bots they never name, through the wildcard group. RFC 9309 is clear on how this works: a crawler must use the group that matches its token, and only "if no matching group exists" does it fall back to the * group.
User-agent: GPTBot
Allow: /
User-agent: *
Disallow: /admin/
User-agent: *
Disallow: /
In our data, 177 files (3.6%) block everyone with the wildcard group. These are often APIs, CDNs, staging hosts or private services that sit high in Tranco because of traffic, not because they publish pages. They raise the "blocked" count for every bot, including Googlebot, which is why Googlebot is blocked on 2.5% of files even though only 11 block it by name.
Legacy and unofficial tokens are still everywhere
Robots.txt files collect old lines. 479 files (9.7%) still name anthropic-ai, a token that does not appear on Anthropic's current crawler page, which lists only ClaudeBot, Claude-User and Claude-SearchBot. The line does no harm, but it does nothing either, and it often sits in place of the tokens that would work.
At the other end, new tokens spread slowly. 303 files name Claude-SearchBot and 495 name OAI-SearchBot, against 839 for GPTBot. xAI-SearchBot, a name reported for Grok search but not documented by xAI, appears in just 3 files.
User-triggered agents are a special case. 287 files block ChatGPT-User in its own group, but OpenAI says that "because these actions are initiated by a user, robots.txt rules may not apply" to it. Perplexity says the same about Perplexity-User. Those lines express a preference, not a control. Our ChatGPT-User guide explains what does work.
Hygiene problems we found
Many files contain errors that change how crawlers behave. None of these are rare edge cases; each one appeared on dozens of top sites.
| Problem | Files or domains | What a crawler does with it |
|---|---|---|
| An HTML page served at /robots.txt (status 200) | 398 domains | Google says it will try to read rules out of HTML and ignore the rest. Most AI crawlers will find no valid rules at all. |
| robots.txt answered with 403 to our study agent | 1,132 domains | RFC 9309 lets a crawler treat a 4xx as "no rules", so a 403 there can mean "crawl everything", the opposite of what the owner wanted. |
| robots.txt answered with a 5xx error | 71 domains | RFC 9309 says to assume complete disallow. Google stops crawling for up to 12 hours, then uses its cached copy. |
| Disallow rules for CSS or JavaScript files | 278 files | Search engines that render pages may not see the page as users do. |
A noindex: line in robots.txt | 21 files | Google stopped supporting it on 1 September 2019. It does nothing. |
| Allow or Disallow before any User-agent line | 26 files | RFC 9309 says crawlers should ignore rules outside a group. |
| A byte order mark at the start of the file | 32 files | Google ignores it. Simpler parsers may misread the first line. |
| Crawl-delay set for Googlebot | 23 files | Google lists crawl-delay as unsupported, so the line is ignored. |
The 403 row needs care. Our study used its own named user agent, and some sites block any non-browser agent at the firewall, including on robots.txt. That is our point: if your firewall hides robots.txt from bots, the bots cannot read your rules. Our guide to why AI bots cannot crawl a site walks through each layer.
New signals: Content-Signal and llms.txt
Two newer ideas are starting to appear in robots.txt files. The first is the Content-Signal line, which Cloudflare adds to sites that turn on its managed robots.txt. It states preferences such as search=yes, ai-train=no. In July 2026 Cloudflare added an optional use value with three levels, immediate, reference and full (Cloudflare, 1 July 2026). We found Content-Signal lines on 125 files (2.5%).
The most common values were ai-train=no, search=yes and the fully open search=yes, ai-input=yes, ai-train=yes. Some files already carry use=reference. Like everything in robots.txt, these are preferences, and the IETF AI Preferences working group is still writing the standard vocabulary for them (IETF aipref charter).
The second idea is llms.txt, a separate file that some sites publish to give AI assistants a summary. It does not belong in robots.txt, but 105 files mention it, usually in a comment. No major AI operator has said it reads llms.txt when deciding what to crawl, so treat it as a helpful extra, not a control.
How sites use Crawl-delay
573 files (11.5%) contain at least one Crawl-delay line. It is not part of RFC 9309, and support differs by crawler: Bing says BingBot honors it, Semrush caps it at 10 seconds, Anthropic supports it for ClaudeBot, and Google ignores it. The values people choose are shown below.
Ten seconds is the most common choice by a wide margin. At one request every ten seconds a crawler can fetch at most 8,640 pages a day, which is fine for a small site and a hard ceiling for a large one. Our AI crawler rate limiting guide covers what to use instead when Crawl-delay is ignored.
How this compares with other studies
Other teams measure AI crawlers from a different angle, and the results fit together. Hostinger looked at real crawler requests across about 5 million hosted sites and found GPTBot's coverage fell from 84% to 12% of sites over 2025, which it put down to "websites blocking AI-training crawlers" (Hostinger, January 2026). Cloudflare's network data showed training crawling growing to about 80% of AI bot activity by July 2025 (Cloudflare, August 2025).
Our study measures stated intent, not traffic. A robots.txt block tells you what the owner asked for. It does not tell you whether the crawler obeyed, and it does not see firewall rules. Cloudflare's 2024 analysis made the same point about its own top-10,000 robots.txt review: these blocks "are reliant on the bot operator respecting robots.txt". Use our numbers to understand policy, and log data to understand behavior.
Limitations you should know about
- The denominator is not "websites". Tranco ranks registered domains by popularity. 2,202 domains in the list did not resolve as a website at all, because they are API, CDN or tracking domains. We only counted the 4,962 that returned a real file.
- One fetch, one agent. We fetched each file once, with our own named agent. A few sites serve different robots.txt files to different agents, and some blocked our agent with a 403. Those sites are counted as "no file", not as "blocks everything".
- robots.txt only. We did not test firewalls, CDN bot rules, meta robots tags or X-Robots-Tag headers. A site can allow GPTBot in robots.txt and still block it at the edge. That is exactly what our AI crawler diagnostic tests for a single page.
- Homepage only. "Blocked" means the token cannot fetch
/. Sites that block only a section, such as/news/, are counted as allowed. - A snapshot. Robots.txt files change. We plan to repeat the fetch and publish the change, keeping this method so the numbers compare.
Check your own robots.txt the same way
You do not need to run a 10,000-domain crawl to see where your site stands. Our free AI crawler checker reads your robots.txt with the same RFC 9309 logic and reports, for each of 248 bots, whether the homepage is allowed, blocked or partly blocked. If you want to test a file before you publish it, paste it into the robots.txt validator. If you want to write one from scratch with a training-versus-search split, use the robots.txt generator.
- Run the checker on your homepage and note which AI search bots are blocked.
- Open your live robots.txt in a browser. Make sure it is plain text, not an HTML page.
- Look for a single block list that mixes training and search tokens. Split it.
- Remove dead lines such as
noindex:andanthropic-aionce the current tokens are in place. - Run the AI crawler diagnostic to see whether a firewall blocks what robots.txt allows.
Download the data and cite this study
The study CSV has one row per user-agent token: how many files name it, how many block it, how many block it in its own group versus the wildcard group, and the block rate in each rank band. You are welcome to reuse it with a link back to this page.
AI Crawler Check (2026). AI crawler rules in the top 10,000 sites' robots.txt files.
Horatos.ai. Data fetched 5 October 2026.
https://aicrawlercheck.com/blog/ai-crawler-robots-txt-study-2026
If you find an error in the data or the method, email the team through the about page and we will correct it and log the change at the top of this article.
Sources and further reading
Every claim above links to the page it came from. These are the primary sources, all read or re-read on 5 October 2026. Where we give a figure from our own data, it comes from our robots.txt study of the top 10,000 sites, and the raw counts are in the downloadable CSV.
- Tranco: a research-oriented top sites ranking
- RFC 9309, Robots Exclusion Protocol (IETF, September 2022)
- Google Search Central: How Google interprets the robots.txt specification
- Google Search Central Blog: A note on unsupported rules in robots.txt (2 July 2019)
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web? (7 April 2026)
- Perplexity: Perplexity crawlers
- Google Search Central: Google's common crawlers (Google-Extended entry)
- Semrush: SemrushBot
- Bing Webmaster Blog: To crawl or not to crawl, that is BingBot's question (May 2012)
- Cloudflare: Content Independence Day, new AI options (1 July 2026)
- Cloudflare: Declaring your AIndependence (3 July 2024)
- Cloudflare: The crawl-to-click gap (29 August 2025)
- Hostinger: 66 billion bot requests analysis (20 January 2026)
- IETF AI Preferences (aipref) working group charter