How We Calculate Your AI Visibility Score

Full transparency on scoring, data sources, and what we can and cannot detect. Last methodology update: August 5, 2026 (v6.0).

Score Composition (100 points)

Bot Access70 pts

Whether the AI crawlers that matter can reach your content, weighted by tier. Tier 1 AI answer engines (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended peers) carry the most weight. Bots that ignore robots.txt by design, deprecated tokens, and opt-out directives are excluded from penalties, because blocking or allowing them in robots.txt changes nothing.

AI Infrastructure15 pts

llms.txt (10 pts) and llms-full.txt (5 pts). Honest framing: no major AI vendor has confirmed consuming these files, and Google's May 2026 AI guidance says they can be ignored. We weight them low. They remain a five-minute, zero-risk hedge, and Lighthouse 13.3 now audits llms.txt.

Technical Readiness15 pts

XML sitemap (6 pts), structured data on the homepage (6 pts), HTTPS (3 pts). These are signals AI crawlers and answer engines demonstrably use for discovery and fact extraction.

The 2026 Bot Taxonomy We Use

Not all AI bots are equal. We classify every user-agent into five behavioral types, because each type needs a different decision from you:

TrainingCollects content to train foundation models (GPTBot, ClaudeBot, Meta-ExternalAgent). Respects robots.txt per vendor documentation.
Search indexBuilds retrieval indexes for AI search (OAI-SearchBot, Claude-SearchBot, PerplexityBot). Blocking these removes you from AI search results.
User-triggeredFetches a page because a human asked (ChatGPT-User, Claude-User, Google-Agent, Gemini Notebook). Google documents that its user-triggered fetchers ignore robots.txt entirely. We mark these "robots.txt not applicable" and never penalize you for them.
Opt-out tokensGoogle-Extended and Applebot-Extended are not crawlers. They are robots.txt control switches that tell an existing crawler not to use your content for AI. We show them in a dedicated panel instead of the bot grid.
UndeclaredAgentic browsers (ChatGPT Atlas, Perplexity Comet, Copilot Actions) and stealth retrieval traffic use normal browser user-agents. No robots.txt rule and no checker tool can detect them. We tell you this instead of pretending otherwise.

How the Check Works

  1. We fetch your robots.txt and parse it per RFC 9309, including longest-match precedence and group merging.
  2. If a WAF or CDN firewall blocks our primary fetch, we retry with fallback strategies and clearly flag results when we cannot confirm the real file.
  3. We evaluate all 196 known user-agent tokens against your rules, including wildcard inheritance.
  4. We check for llms.txt, llms-full.txt, your XML sitemap, homepage structured data, HTTPS, and the IETF Web Bot Auth directory at /.well-known/http-message-signatures-directory.
  5. We scan your robots.txt for deprecated tokens (anthropic-ai, claude-web, google-notebooklm) that vendors have retired, a common source of false security.

What We Cannot Detect (And Nobody Can)

Data Sources

Our bot database is built from primary vendor documentation and verified security research, refreshed continuously. Intelligence version: 2026-08-05.

Ready to check your site?

Run a free scan against all 196 bots with this exact methodology.

Check My Website