#1 AI Crawler Checker
Is Your Website Invisible to
ChatGPT, Gemini & other AI Crawlers?
Most sites block AI crawlers by accident.
Enter your URL to see exactly which of
196 bots can reach you.
Try:
100% free No sign-up Results in seconds
Analyzing website...
Crawler access results
- -
- Bot Access
- /70
- -
- AI Infra
- /15
- -
- Tech Ready
- /15
Category Overview
Access status summary across all 8 bot categories
Bot Access Status
Each bot is checked against robots.txt rules (the industry standard for bot access policy). Results are based on what the site owner explicitly allows or blocks.
AI Infrastructure Files
These files tell AI engines what your site is about and how to read it. Honest note: no major AI vendor has confirmed reading llms.txt yet, but it is a five-minute, zero-risk hedge that Lighthouse now audits.
Fix This First: Priority Action Plan
Ranked by score impact. Start at number 1 and work down.
Export Full Report
Download your AI Visibility report - branded & verified by Horatos.ai
Embed Your Score Badge
Show your AI Visibility Score on your website. The badge links back to your live results.
Share your results
What invisibility costs you
AI assistants are where buyers now start their research. When they cannot read your site, the loss never shows up in your analytics, because the visit never happens.
-
01
Silent loss
No warning, no error, no dip to investigate
Nothing breaks when GPTBot or ClaudeBot is turned away. Your rankings look fine, your dashboards stay green, and you quietly stop appearing in AI answers. There is no alert for this in Search Console, in your analytics, or anywhere else, so most site owners find out months late, or never find out at all.
-
02
Competitor gain
The citation goes to whoever is readable
Every question ChatGPT or Perplexity answers needs a source, and the model picks from the pages it is allowed to read. If yours are blocked, the answer still gets written, it just names a competitor whose robots.txt let the crawler in. You are not ranked lower. You are not in the running.
-
03
Compounding gap
AI visibility compounds, like search did
The sites being cited today become the default sources models keep returning to, and being cited makes a page more likely to be cited again. Every month you stay blocked is a month of citations, mentions and trust accruing to someone else, and buying it back later costs far more than not losing it.
What is the best project management tool for a growing team?
Blocked
Your site cannot be read
For a growing team, the options people rate most highly are Competitor A and Competitor B. Both scale well past twenty seats and integrate with the tools you already use…
Sources: competitor-a.com, competitor-b.com
Readable
Your site can be read
For a growing team, your product is the one most often recommended: it is built for teams at exactly this stage, and its own documentation sets out the pricing and the seat limits…
Sources: yoursite.com, competitor-a.com
How AI Crawler Check works
Three steps, about ten seconds. Nothing to install, nothing to sign up for, and no change to your website: paste your address and read the answer.
-
You do this
Paste your web address
We visit your site the same way an AI assistant would and read the small text files that tell visiting robots what they may look at. We read a handful of files, not your whole website, and we do not save the address you typed.
0fields to fill in. No account, no email, no card. -
We do this
We try every AI robot by name
Each robot has its own name, and your rules can allow one and block another without you realising. We test all 196 of them, one at a time, using the same matching order the real robots use. That order is where reading a rules file by eye usually goes wrong.
196robots tested individually, across 8 groups. -
You get this
A score, and the exact fix
A single number out of 100, a plain list of which assistants cannot see you and why, and ready-made text you can copy and paste to fix it. The fixes are ordered by how much each one is worth, so the first one you apply is the one that matters most.
~10sfrom pasting the address to reading the result.
What your free AI crawler report shows
No sign-up wall between you and the answer. Every report has the same four parts: a score, the robots that cannot reach you, the reason for each, and the fix.
Sample result
yoursite.comHigh risk: invisible to the assistants that matter
3 of the 5 most important AI crawlers are turned away by your robots.txt, and your llms.txt is missing.
- 21/70Crawler access
- 6/15AI files
- 15/15Technical
- GPTBot · OpenAI robots.txt line 4 · Disallow: / Blocked
- ClaudeBot · Anthropic robots.txt line 4 · Disallow: / Blocked
- PerplexityBot · Perplexity X-Robots-Tag header · noai Blocked
- Google-Extended · Google Gemini No rule matched · access permitted Allowed
The full report lists all 196 robots this way, grouped by how much each one is worth to you.
Your fix
copy & paste- +28 ptsRemove the blanket Disallow on line 4 and allow the three named crawlers above.
- +10 ptsAdd an llms.txt file at your root so assistants can find your key pages without guessing.
- +6 ptsDrop the noai header your CDN is adding to every response.
Every snippet is generated for your site, not a template, and one click copies it.
AI Crawler Check vs typical AI bot checkers
Most checkers were built before AI search mattered and test a thin slice of the crawler landscape. Here is the difference, row by row.
| Capability | AI Crawler Check | Typical AI-only tools |
|---|---|---|
| Bots tested | 196 bots | 30 to 40 |
| Categories covered | 8 categories | 1 to 3 (AI only) |
| Agentic crawlers (Google-Agent, OAI-AdsBot) | ||
| Unified GEO score (58 checks, 0 to 100) | ||
| Signals read per URL | robots.txt, meta, headers, llms.txt, sitemap, schema | robots.txt only |
| Shows the rule that caused each result | ||
| Flags renamed and retired robot tokens | ||
| Many URLs in one run | Up to 20 | One at a time |
| llms.txt detection with strict validation | Partial | |
| Honest WAF handling (15+ providers) | ||
| Free tools included | 6 free tools | Rare |
| Sign-up required | Never | Often |
The largest AI crawler database on the open web
Every robot we test has a page of its own: who runs it, what it takes, whether it obeys robots.txt, and the exact line to allow or block it. All of it free to read, no sign-up.
- 71AI and LLM bots The robots that feed answer engines: GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest. Browse AI bots
- 48Scrapers Bulk harvesters and dataset builders such as Bytespider and CCBot, with the rule to stop each one. Browse scrapers
- 24Search engines Bingbot, DuckDuckBot, Yandex and the regional engines that still send real clicks to your pages. Browse search engines
- 21SEO tools AhrefsBot, SemrushBot, Screaming Frog and the other auditors that crawl you on someone else's behalf. Browse SEO tools
- 16Google bots Googlebot, Google-Extended, AdsBot and Google-Agent, each with a different job and a different rule. Browse Google bots
- 7Social bots Link unfurlers and preview fetchers. Block one and your links stop showing a title and an image. Browse social bots
- 5Other agents Agentic browsers and assistants that click and type on a real person's behalf rather than indexing. Browse other agents
- 4Cloud services Platform and CDN fetchers. Harmless until one is caught by a blanket rule meant for someone else. Browse cloud services
- 91Operator profiles Every company running a crawler, which robots it owns, and how its tokens have been renamed over time. Browse operators
- 20Head-to-head comparisons GPTBot against ClaudeBot, Google-Extended against Googlebot: what each pair actually does differently. Compare crawlers
- 146Research articles Per-robot guides, llms.txt standards work, and analysis of every new crawler the week it appears. Read the research
- 58GEO checks per scan The full checklist every URL runs: robots.txt, meta robots, X-Robots-Tag, llms.txt, sitemap and schema. See the methodology
Latest from the crawler database
New robots appear every few weeks and old ones get renamed without notice. This is what changed most recently, written up the week it happened.
-
Google-Agent (Project Mariner): the AI bot that browses like a human
Google's first agent that clicks, types and fills forms on a real person's behalf. Robots.txt does not stop it.
-
WebMCP explained: how to make your website AI agent-ready
The emerging standard that lets an agent use your site's features directly instead of guessing at your HTML.
-
What is ChatGPT-User? OpenAI's search crawler explained
The robot that fetches your page live when someone asks ChatGPT a question. Blocking it removes you from the answer.
-
What is Applebot-Extended? The Apple Intelligence crawler
Apple split training access away from search indexing. One token now decides whether Siri can quote you.
-
ClaudeBot and Anthropic's AI crawlers: the complete guide
Anthropic runs more than one robot and they do different jobs. Which token you write decides which one you stop.
-
PerplexityBot: how Perplexity AI crawls the web
Perplexity cites its sources by name in every answer, which makes access to it unusually valuable.
Six free tools to fix what the scan finds
A report is only useful if you can act on it. These write the actual files for you, check the ones you already have, and none of them ask for an account.
-
Most complete
GEO Site Audit
The full 58-point GEO and AEO audit: crawler access, structure, schema and the AI files, scored out of 100 with every fix written out.
Run the audit -
6 presets
robots.txt Generator
Choose which robots may read you, one by one, and get a correct file back. Six ready-made presets from fully open to search only.
Build a file -
Catches typos
robots.txt Validator
Paste your existing file and see the errors that matter: bad syntax, rules in the wrong order, and robot names that no longer exist.
Check my file -
New standard
llms.txt Generator
The emerging standard that tells an assistant which of your pages matter and what each one covers, so it stops guessing.
Generate llms.txt -
JSON-LD
Schema Generator
Structured data an answer engine can read without interpreting your layout. Fill in the fields, copy the block into your page head.
Generate schema -
Up to 20 URLs
Batch Checker
Test many pages in one run. Useful when a template is the thing that is blocking you rather than any single page.
Check in bulk
Six ways a site goes invisible by accident
Almost nobody decides to hide from AI assistants. It happens through a line copied from a forum, a plugin default, or a header a hosting provider adds for you. Here is what we find most often, and what each one does to your score.
-
The blanket disallow
Two lines meant to keep one scraper out. They apply to every robot ever written, including the ones that would have quoted you.
User-agent: *
Disallow: / -
A header you never wrote
Some hosts and security plugins add an opt-out header to every response by default. Your robots.txt says yes and your server says no.
X-Robots-Tag: noai, noimageai
-
A name that no longer exists
Operators rename their robots. A rule written for the old token stops matching anything on the day of the change, silently, and nothing warns you.
User-agent: anthropic-ai
Disallow: / -
A firewall that answers first
Bot protection can challenge an AI crawler before your rules are ever read. The robot sees a security page, not your content, and moves on.
403 Forbidden · managed bot rule
-
Content that arrives empty
If your text is drawn by JavaScript after the page loads, a robot that does not run scripts receives a shell with nothing in it to quote.
<div id="root"></div>
-
Nothing to guide the reader
Access alone is not readability. With no llms.txt, no schema and no clear headings, an assistant has to guess which of your pages answers the question.
GET /llms.txt → 404
Frequently Asked Questions
What does an AI crawler checker actually do?
It fetches your robots.txt, meta robots tags, X-Robots-Tag headers and AI files such as llms.txt, then evaluates each rule against the user-agent string of every known AI crawler. The result tells you which AI systems are allowed to read your content and which are blocked, either intentionally or by accident.
Why would blocking AI crawlers hurt my business?
When GPTBot, Google-Extended, ClaudeBot or PerplexityBot cannot read your pages, AI assistants cannot cite you as a source. Your competitors get recommended instead. Unlike a search ranking drop, this failure is silent: no error, no traffic alert, no warning in Search Console.
Is this tool really free, and do I need an account?
Yes, completely free with no account, no email and no credit card. All 196 bot checks, the 58-point GEO audit and all six free tools are available immediately. We do not store the URLs you test.
What is the difference between SEO and GEO?
SEO optimises for ranked links in a results page. GEO, or Generative Engine Optimization, optimises for being quoted inside an AI-generated answer. GEO depends on crawler access, machine-readable structure, schema markup and clear factual statements, which is what this tool measures.
How is the AI Visibility Score calculated?
The score weights crawler access at 70 percent, AI infrastructure files at 15 percent and technical readiness at 15 percent. Bots are weighted by real-world importance, so blocking GPTBot costs more than blocking a minor scraper. The full formula is documented on our methodology page.
Should I allow or block AI crawlers on my website?
It depends on the robot, which is why a single blanket rule is almost always the wrong answer. Retrieval crawlers such as ChatGPT-User and PerplexityBot fetch your page to answer a live question and cite you by name, so blocking them removes you from answers your customers are already asking for. Training-only crawlers such as CCBot copy your content into a dataset with no link back, and blocking those costs you nothing. Bulk scrapers that ignore robots.txt entirely need a firewall rule rather than a robots.txt line. The practical position for most businesses is to allow the crawlers that cite, block the ones that only harvest, and decide the training-plus-retrieval ones like GPTBot on your own commercial terms.
How do I know if ChatGPT can read my website?
Four things have to line up, and checking only the first is the usual mistake. Your robots.txt must not disallow GPTBot or ChatGPT-User. Your pages must not carry a noai or noindex instruction in a meta tag or an X-Robots-Tag response header, which your host or a plugin may be adding without your knowledge. Your firewall must not challenge the crawler before it reaches your content. And your text has to be present in the HTML the server returns, not drawn in afterwards by JavaScript. Asking ChatGPT whether it can see your site does not test any of this reliably, because the model may answer from memory. Running the URL through this checker tests all four in one pass.
How often should I run an AI crawler check?
Once a quarter as a baseline, and immediately after any change to your hosting, CDN, security settings or CMS plugins, because those are the changes that alter robots.txt and response headers without anyone editing a file on purpose. Also check when a new crawler is announced: operators add and rename robots several times a year, and a rule written for a token that no longer exists stops working the day the name changes, with no error and no notification. The scan is free and takes about ten seconds, so there is no reason to wait for a reason.
A free tool by
Horatos.ai
AI-native growth marketing agency in Singapore
We built this checker because we needed it ourselves. If you want the strategy behind the score, that is the work we do for clients.