How many of the web's biggest sites block AI crawlers in 2026, and which ones? We fetched the robots.txt file of every domain in the Tranco top 10,000 on 2026-10-05 and parsed 4,962 valid files the way RFC 9309 says a crawler must. This page shows what we found, how we did it, and where the method has limits. The raw counts are free to download as a CSV.

What changed in this update
  • : First published. Data fetched the same day.

Key Takeaways

  • 14.3% of sites block GPTBot, the most named AI crawler. CCBot (15.4%) and Bytespider (15%) are blocked slightly more often.
  • 19.2% block at least one AI training crawler, but only 10.8% block an AI search crawler.
  • About half of GPTBot blockers keep OpenAI search open. 347 of 710 still allow OAI-SearchBot.
  • 79.9% of Google-Extended blockers keep Googlebot, which is exactly how Google designed the token.
  • The top 1,000 sites block AI crawlers about 1.5 times as often as sites ranked 5,001 to 10,000.
  • 398 sites answered with an HTML page instead of robots.txt rules, usually a challenge or error page.
How AI crawlers work: pipeline from website to AI answer Flow diagram showing four stages: your website is checked against robots.txt rules, an AI crawler reads allowed content, the content feeds an AI model, and the model produces AI answers that can cite your site. Your Website pages + content example.com robots.txt User-agent: GPTBot Allow: / the gatekeeper AI Crawler GPTBot, ClaudeBot, PerplexityBot... reads allowed pages AI Model training + retrieval AI Answers citing your site Block the crawler at step 2 and your content never reaches the AI answer at step 4.
How AI crawlers work: from your website through robots.txt to AI-generated answers.

Key findings at a glance

Here are the headline numbers. Every percentage uses the same denominator, the 4,962 domains that returned a real robots.txt file, unless the row says otherwise. We explain the denominator in detail in the limitations section, because it changes how you should read the rest of the page.

MeasureCountShare
Domains in the list10,000100%
Domains that returned any HTTP response7,31573.2%
Valid robots.txt files parsed4,96249.6% of the list
Block at least one AI training crawler95219.2%
Block at least one AI search crawler53410.8%
Block every crawler with User-agent: * and Disallow: /1773.6%
List a sitemap3,34167.3%
Use Crawl-delay anywhere57311.5%
Use a Content-Signal line1252.5%

"Training crawler" here means at least one of GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Meta-ExternalAgent or Applebot-Extended. "AI search crawler" means at least one of OAI-SearchBot, Claude-SearchBot or PerplexityBot. We chose these groups because each operator's own documentation describes the token that way, as we explain below.

Which AI crawlers are blocked most often?

The chart counts a token as blocked when the rules that apply to it stop it from fetching the homepage. That can happen in two ways: the file names the token in its own group, or the file has no group for it and the wildcard group blocks everyone. We show both, because the difference matters when you write your own file.

AI crawlers blocked from the whole siteShare of 4,962 valid robots.txt files from the Tranco top 10,000 that block each token from the whole site (fetched 2026-10-05).AI crawlers blocked from the whole siteCCBot15.4%Bytespider15%GPTBot14.3%ClaudeBot13.4%Google-Extended12%Meta-ExternalAgent12%Applebot-Extended11.2%Amazonbot11%PerplexityBot10.5%ChatGPT-User9.1%OAI-SearchBot7.4%Claude-SearchBot7.3%GrokBot4.8%Source: AI Crawler Check robots.txt study, October 2026. n = 4,962 files.
Share of 4,962 valid robots.txt files from the Tranco top 10,000 that block each token from the whole site (fetched 2026-10-05).
TokenNamed in the fileBlocked (any group)Blocked by its own groupBlocked only by *
CCBot718 (14.5%)766 (15.4%)588178
Bytespider638 (12.9%)746 (15%)565181
GPTBot839 (16.9%)710 (14.3%)549161
ClaudeBot710 (14.3%)665 (13.4%)494171
Google-Extended655 (13.2%)597 (12%)431166
Meta-ExternalAgent506 (10.2%)593 (12%)415178
Applebot-Extended474 (9.6%)557 (11.2%)383174
Amazonbot475 (9.6%)545 (11%)370175
PerplexityBot613 (12.4%)520 (10.5%)356164
ChatGPT-User566 (11.4%)450 (9.1%)287163
OAI-SearchBot495 (10%)367 (7.4%)205162
Claude-SearchBot303 (6.1%)361 (7.3%)186175
GrokBot47 (0.9%)236 (4.8%)43193

Three things stand out. First, the old training crawlers sit at the top: Common Crawl's CCBot and ByteDance's Bytespider were the first bots many site owners heard about, and their blocks have stayed. Second, GPTBot is the token people write most often (16.9% of files name it), yet it is not the most blocked, because some files name it only to allow it. Third, the search crawlers in teal are blocked about half as often as the training crawlers from the same companies.

Why GrokBot looks high

Only 47 files name GrokBot, but 236 block it. The difference, 193 files, is sites that block every crawler with the wildcard group. xAI has not documented a crawler, so few owners write a rule for it. See our GrokBot profile for what is and is not known.

OpenAI, Anthropic and Google each publish separate controls for training and for search. OpenAI's crawler page says each setting "is independent of the others", so a site can allow OAI-SearchBot to appear in ChatGPT search while disallowing GPTBot for training. Anthropic's help page describes ClaudeBot for training and Claude-SearchBot for search. Google's crawler list says Google-Extended "does not impact a site's inclusion in Google Search".

So we asked a simple question: when a site blocks the training token, does it keep the search token open? The answer is "about half the time" for OpenAI and Anthropic, and "almost always" for Google.

CompanySites blocking the training tokenOf those, search token still allowedShare that split the two
OpenAI (GPTBot vs OAI-SearchBot)71034748.9%
Anthropic (ClaudeBot vs Claude-SearchBot)66530746.2%
Google (Google-Extended vs Googlebot)59747779.9%

The Google number is high because almost nobody blocks Googlebot on purpose: only 11 files block it in its own group. The OpenAI and Anthropic numbers are the interesting ones. Roughly half of the sites that opted out of training also opted out of AI search, often with one copy-pasted block list. If your goal was "no training, but please cite me", a block list like that works against you. Our robots.txt best practices guide shows the split template.

Anatomy of a robots.txt file with annotations An annotated robots.txt example. The User-agent line targets a specific bot such as GPTBot. Disallow blocks paths, Allow grants exceptions, the wildcard user-agent covers every other bot, and the Sitemap line points crawlers to your XML sitemap. # AI crawler rules User-agent: GPTBot Disallow: /private/ Allow: /blog/ User-agent: * Sitemap: /sitemap.xml Targets one bot by name Each bot reads only its own section Blocks specific paths Grants exceptions Allow overrides broader Disallow Wildcard = every other bot Helps crawlers find pages Always declare your sitemap
Anatomy of a robots.txt file: user-agent targeting, allow and disallow rules, and sitemap declaration.

Bigger sites block AI crawlers more often

We split the list into three bands by Tranco rank. Each band has its own denominator: 505 valid files in the top 1,000, 1,625 in ranks 1,001 to 5,000, and 2,832 in ranks 5,001 to 10,000.

TokenTop 1,000Rank 1,001 to 5,000Rank 5,001 to 10,000
GPTBot20.8%14%13.3%
ClaudeBot20.2%13.3%12.3%
Google-Extended18.8%11.8%11%
CCBot23%14.8%14.5%
Bytespider23.8%14.3%13.9%
OAI-SearchBot12.1%7.1%6.7%
PerplexityBot17.4%10.2%9.4%
Googlebot3.4%2.6%2.3%

For every AI token, the top band blocks the most. GPTBot goes from 13.3% in the lowest band to 20.8% in the top 1,000. Cloudflare saw the same pattern in 2024: "the higher-ranked (more popular) an Internet property is, the more likely it is to be targeted by AI bots, and correspondingly, the more likely it is to block such requests" (Cloudflare, July 2024). Big publishers have content licensing deals to protect and legal teams who read the terms. Smaller sites more often want the visibility.

Named is not the same as blocked

A common mistake when reading robots.txt surveys is to count every mention of a bot as a block. Many files name a crawler only to allow it, or to give it a crawl delay. And many files block bots they never name, through the wildcard group. RFC 9309 is clear on how this works: a crawler must use the group that matches its token, and only "if no matching group exists" does it fall back to the * group.

A file that names GPTBot but does not block it
User-agent: GPTBot
Allow: /

User-agent: *
Disallow: /admin/
A file that blocks GPTBot without naming it
User-agent: *
Disallow: /

In our data, 177 files (3.6%) block everyone with the wildcard group. These are often APIs, CDNs, staging hosts or private services that sit high in Tranco because of traffic, not because they publish pages. They raise the "blocked" count for every bot, including Googlebot, which is why Googlebot is blocked on 2.5% of files even though only 11 block it by name.

Legacy and unofficial tokens are still everywhere

Robots.txt files collect old lines. 479 files (9.7%) still name anthropic-ai, a token that does not appear on Anthropic's current crawler page, which lists only ClaudeBot, Claude-User and Claude-SearchBot. The line does no harm, but it does nothing either, and it often sits in place of the tokens that would work.

At the other end, new tokens spread slowly. 303 files name Claude-SearchBot and 495 name OAI-SearchBot, against 839 for GPTBot. xAI-SearchBot, a name reported for Grok search but not documented by xAI, appears in just 3 files.

User-triggered agents are a special case. 287 files block ChatGPT-User in its own group, but OpenAI says that "because these actions are initiated by a user, robots.txt rules may not apply" to it. Perplexity says the same about Perplexity-User. Those lines express a preference, not a control. Our ChatGPT-User guide explains what does work.

Hygiene problems we found

Many files contain errors that change how crawlers behave. None of these are rare edge cases; each one appeared on dozens of top sites.

ProblemFiles or domainsWhat a crawler does with it
An HTML page served at /robots.txt (status 200)398 domainsGoogle says it will try to read rules out of HTML and ignore the rest. Most AI crawlers will find no valid rules at all.
robots.txt answered with 403 to our study agent1,132 domainsRFC 9309 lets a crawler treat a 4xx as "no rules", so a 403 there can mean "crawl everything", the opposite of what the owner wanted.
robots.txt answered with a 5xx error71 domainsRFC 9309 says to assume complete disallow. Google stops crawling for up to 12 hours, then uses its cached copy.
Disallow rules for CSS or JavaScript files278 filesSearch engines that render pages may not see the page as users do.
A noindex: line in robots.txt21 filesGoogle stopped supporting it on 1 September 2019. It does nothing.
Allow or Disallow before any User-agent line26 filesRFC 9309 says crawlers should ignore rules outside a group.
A byte order mark at the start of the file32 filesGoogle ignores it. Simpler parsers may misread the first line.
Crawl-delay set for Googlebot23 filesGoogle lists crawl-delay as unsupported, so the line is ignored.

The 403 row needs care. Our study used its own named user agent, and some sites block any non-browser agent at the firewall, including on robots.txt. That is our point: if your firewall hides robots.txt from bots, the bots cannot read your rules. Our guide to why AI bots cannot crawl a site walks through each layer.

New signals: Content-Signal and llms.txt

Two newer ideas are starting to appear in robots.txt files. The first is the Content-Signal line, which Cloudflare adds to sites that turn on its managed robots.txt. It states preferences such as search=yes, ai-train=no. In July 2026 Cloudflare added an optional use value with three levels, immediate, reference and full (Cloudflare, 1 July 2026). We found Content-Signal lines on 125 files (2.5%).

The most common values were ai-train=no, search=yes and the fully open search=yes, ai-input=yes, ai-train=yes. Some files already carry use=reference. Like everything in robots.txt, these are preferences, and the IETF AI Preferences working group is still writing the standard vocabulary for them (IETF aipref charter).

The second idea is llms.txt, a separate file that some sites publish to give AI assistants a summary. It does not belong in robots.txt, but 105 files mention it, usually in a comment. No major AI operator has said it reads llms.txt when deciding what to crawl, so treat it as a helpful extra, not a control.

How sites use Crawl-delay

573 files (11.5%) contain at least one Crawl-delay line. It is not part of RFC 9309, and support differs by crawler: Bing says BingBot honors it, Semrush caps it at 10 seconds, Anthropic supports it for ClaudeBot, and Google ignores it. The values people choose are shown below.

Most common Crawl-delay values (seconds)Number of Crawl-delay lines with each value, across 4,962 robots.txt files from the Tranco top 10,000.Most common Crawl-delay values (seconds)10 s3651 s2715 s2322 s13830 s10360 s5020 s423 s4115 s39Source: AI Crawler Check robots.txt study, October 2026.
Number of Crawl-delay lines with each value, across 4,962 robots.txt files from the Tranco top 10,000.

Ten seconds is the most common choice by a wide margin. At one request every ten seconds a crawler can fetch at most 8,640 pages a day, which is fine for a small site and a hard ceiling for a large one. Our AI crawler rate limiting guide covers what to use instead when Crawl-delay is ignored.

How this compares with other studies

Other teams measure AI crawlers from a different angle, and the results fit together. Hostinger looked at real crawler requests across about 5 million hosted sites and found GPTBot's coverage fell from 84% to 12% of sites over 2025, which it put down to "websites blocking AI-training crawlers" (Hostinger, January 2026). Cloudflare's network data showed training crawling growing to about 80% of AI bot activity by July 2025 (Cloudflare, August 2025).

Our study measures stated intent, not traffic. A robots.txt block tells you what the owner asked for. It does not tell you whether the crawler obeyed, and it does not see firewall rules. Cloudflare's 2024 analysis made the same point about its own top-10,000 robots.txt review: these blocks "are reliant on the bot operator respecting robots.txt". Use our numbers to understand policy, and log data to understand behavior.

Limitations you should know about

  • The denominator is not "websites". Tranco ranks registered domains by popularity. 2,202 domains in the list did not resolve as a website at all, because they are API, CDN or tracking domains. We only counted the 4,962 that returned a real file.
  • One fetch, one agent. We fetched each file once, with our own named agent. A few sites serve different robots.txt files to different agents, and some blocked our agent with a 403. Those sites are counted as "no file", not as "blocks everything".
  • robots.txt only. We did not test firewalls, CDN bot rules, meta robots tags or X-Robots-Tag headers. A site can allow GPTBot in robots.txt and still block it at the edge. That is exactly what our AI crawler diagnostic tests for a single page.
  • Homepage only. "Blocked" means the token cannot fetch /. Sites that block only a section, such as /news/, are counted as allowed.
  • A snapshot. Robots.txt files change. We plan to repeat the fetch and publish the change, keeping this method so the numbers compare.

Check your own robots.txt the same way

You do not need to run a 10,000-domain crawl to see where your site stands. Our free AI crawler checker reads your robots.txt with the same RFC 9309 logic and reports, for each of 248 bots, whether the homepage is allowed, blocked or partly blocked. If you want to test a file before you publish it, paste it into the robots.txt validator. If you want to write one from scratch with a training-versus-search split, use the robots.txt generator.

  1. Run the checker on your homepage and note which AI search bots are blocked.
  2. Open your live robots.txt in a browser. Make sure it is plain text, not an HTML page.
  3. Look for a single block list that mixes training and search tokens. Split it.
  4. Remove dead lines such as noindex: and anthropic-ai once the current tokens are in place.
  5. Run the AI crawler diagnostic to see whether a firewall blocks what robots.txt allows.

Download the data and cite this study

The study CSV has one row per user-agent token: how many files name it, how many block it, how many block it in its own group versus the wildcard group, and the block rate in each rank band. You are welcome to reuse it with a link back to this page.

Suggested citation
AI Crawler Check (2026). AI crawler rules in the top 10,000 sites' robots.txt files.
Horatos.ai. Data fetched 5 October 2026.
https://aicrawlercheck.com/blog/ai-crawler-robots-txt-study-2026

If you find an error in the data or the method, email the team through the about page and we will correct it and log the change at the top of this article.

Sources and further reading

Every claim above links to the page it came from. These are the primary sources, all read or re-read on 5 October 2026. Where we give a figure from our own data, it comes from our robots.txt study of the top 10,000 sites, and the raw counts are in the downloadable CSV.

  1. Tranco: a research-oriented top sites ranking
  2. RFC 9309, Robots Exclusion Protocol (IETF, September 2022)
  3. Google Search Central: How Google interprets the robots.txt specification
  4. Google Search Central Blog: A note on unsupported rules in robots.txt (2 July 2019)
  5. OpenAI: Overview of OpenAI crawlers
  6. Anthropic: Does Anthropic crawl data from the web? (7 April 2026)
  7. Perplexity: Perplexity crawlers
  8. Google Search Central: Google's common crawlers (Google-Extended entry)
  9. Semrush: SemrushBot
  10. Bing Webmaster Blog: To crawl or not to crawl, that is BingBot's question (May 2012)
  11. Cloudflare: Content Independence Day, new AI options (1 July 2026)
  12. Cloudflare: Declaring your AIndependence (3 July 2024)
  13. Cloudflare: The crawl-to-click gap (29 August 2025)
  14. Hostinger: 66 billion bot requests analysis (20 January 2026)
  15. IETF AI Preferences (aipref) working group charter