How We Calculate Your AI Visibility Score
Full transparency on scoring, data sources, and what we can and cannot detect. Last methodology update: August 5, 2026 (v6.0).
Score Composition (100 points)
Whether the AI crawlers that matter can reach your content, weighted by tier. Tier 1 AI answer engines (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended peers) carry the most weight. Bots that ignore robots.txt by design, deprecated tokens, and opt-out directives are excluded from penalties, because blocking or allowing them in robots.txt changes nothing.
llms.txt (10 pts) and llms-full.txt (5 pts). Honest framing: no major AI vendor has confirmed consuming these files, and Google's May 2026 AI guidance says they can be ignored. We weight them low. They remain a five-minute, zero-risk hedge, and Lighthouse 13.3 now audits llms.txt.
XML sitemap (6 pts), structured data on the homepage (6 pts), HTTPS (3 pts). These are signals AI crawlers and answer engines demonstrably use for discovery and fact extraction.
The 2026 Bot Taxonomy We Use
Not all AI bots are equal. We classify every user-agent into five behavioral types, because each type needs a different decision from you:
How the Check Works
- We fetch your robots.txt and parse it per RFC 9309, including longest-match precedence and group merging.
- If a WAF or CDN firewall blocks our primary fetch, we retry with fallback strategies and clearly flag results when we cannot confirm the real file.
- We evaluate all 196 known user-agent tokens against your rules, including wildcard inheritance.
- We check for llms.txt, llms-full.txt, your XML sitemap, homepage structured data, HTTPS, and the IETF Web Bot Auth directory at /.well-known/http-message-signatures-directory.
- We scan your robots.txt for deprecated tokens (anthropic-ai, claude-web, google-notebooklm) that vendors have retired, a common source of false security.
What We Cannot Detect (And Nobody Can)
- Whether a bot actually honors your robots.txt. We report documented vendor behavior and known violations (for example, Cloudflare's August 2025 findings on Perplexity stealth crawling).
- Agentic browser traffic that presents a standard Chrome user-agent.
- Server-side blocks (WAF rules, IP blocks) that do not appear in robots.txt. Your robots.txt may allow a bot your firewall silently rejects.
- Whether AI engines actually cite you. Access is necessary but not sufficient for citations.
Data Sources
Our bot database is built from primary vendor documentation and verified security research, refreshed continuously. Intelligence version: 2026-08-05.
- OpenAI, Anthropic, Google, Microsoft, Apple, Meta, Perplexity official crawler documentation
- Published IP ranges and verification endpoints where vendors provide them
- Cloudflare Radar verified bots data and security research
- IETF Web Bot Auth draft specifications
Ready to check your site?
Run a free scan against all 196 bots with this exact methodology.
Check My Website