Each AI engine crawls, ranks, and cites a little differently. Optimizing for one is good. Understanding the differences across all of them is how you win.
In this guide you will learn how to become a Perplexity source. We will keep it practical, with clear steps, visual breakdowns, and specific actions you can take today. The first step in any AI visibility project is to run a free AI crawler check on your website so you know exactly where you stand against the 196 bots we track across 8 categories.
Key Takeaways
- GEO and SEO share most of their foundation, but the differences decide AI citations.
- Citability, authority, and crawl access are the three pillars that matter most.
- You can measure progress with an AI Visibility Score and AI referral tracking.
- Start with a free baseline using the free AI crawler check.
The Real Shape of Perplexity Citations
Perplexity is the one engine in this cluster where the documented behaviour and the observed behaviour do not agree, and that disagreement has to be the subject of the post rather than a caveat at the end of it. Every other engine here can be discussed in terms of which token to set. This one cannot, because the value of setting the token is exactly what is in question.
The database records the Perplexity crawler as inconsistent on robots.txt, with the note that the documentation says it respects robots.txt but that Cloudflare published evidence in August 2025 of undeclared stealth crawlers bypassing robots.txt, and de-listed Perplexity from its verified bots programme. The associated user agent is recorded as named in the same investigation. The database also carries the explicit instruction not to rely on robots.txt alone for a hard block. So there are two separate jobs here: getting cited, which is mostly content work, and preventing access, which cannot be done cooperatively.
Four Perplexity Policy Choices Worth Making Deliberately
Decide which of the two jobs you are doing, because the answers point in opposite directions
If you want Perplexity citations, this is a content and structure problem and the access question barely arises. If you want to prevent Perplexity access, the recorded inconsistency means a cooperative control is not sufficient and you need an edge or server-side rule. Most confusion in this area comes from mixing the two conversations in one paragraph. Write down which one you are having before you touch a file.
If you are blocking, do it where compliance is not optional
The database is unusually direct in saying not to rely on robots.txt alone for a hard block here. That means the deny rule belongs at the edge, where enforcement does not depend on the client choosing to cooperate. Keep the robots.txt rule anyway, because it states the intent and correctly handles the agents that do comply, but do not count it as the control. The distinction between a request and a rule is the whole point.
If you are pursuing citations, note that this engine surfaces sources prominently
Perplexity foregrounds its sources, which changes the payoff of being citable rather than merely crawled. A page that answers one question cleanly, in a passage that stands alone, is much more likely to be quoted than a comprehensive page that answers twelve questions in passing. The practical implication is to prefer a specific page over a broad one when you have the choice, because specificity is what makes a passage safe to lift.
Record verification status rather than trusting the header
Because undeclared crawlers are documented here, the user agent is a particularly weak signal for this operator. The database does record published address ranges for the Perplexity agents, so a forward-confirmed check is possible for the declared traffic. Do it, and store the verdict as its own field. Traffic that claims this agent and fails the check is the interesting number, and you cannot produce it at all if your logs only record what the request said about itself.
The Failure Mode to Watch For in Perplexity Citations
Reporting a robots.txt rule as a block when the database records the agent as inconsistent
The failure is not the rule, it is the confidence. Someone adds a disallow, ticks the item off, and reports the site as blocked to whoever asked. The wording in the database is that documentation and published evidence disagree and that robots.txt alone should not be relied on for a hard block. So the honest report is that a request has been made and partly honoured, which is a different sentence with different consequences for anyone downstream deciding whether sensitive content is safe to publish. Say cooperative or enforced. Never say blocked without specifying which.
What to Confirm Before You Trust Your Perplexity Policy
The check that matters here: For the traffic claiming this agent, produce three numbers: confirmed against published ranges, contradicted, and not checkable. If you can only produce one number, you are reporting claims rather than access.
Where to Go From Here
Perplexity Policy is one piece of a larger picture. The full list of AI crawlers documents every crawler we track with its operator, purpose and safety rating, and the batch checker audits many sites in one pass if you manage a portfolio.
Turn the guidance above into a concrete change, then confirm it worked. An AI bot access checker shows you exactly which of the 196 bots can reach your content today.
Your Perplexity Policy Action Checklist
Five concrete steps, specific to what this guide covered. Work through them in order, changing one thing at a time so you can tell which change produced the result.
- □Establish a baseline with an AI bot access checker and write down the score before you change anything.
- □Apply the single highest-impact change from this guide, on its own, so you can attribute the result.
- □Validate the change with the robot checker before it reaches production.
- □Re-measure and compare against your baseline rather than against expectation.
- □Schedule a recurring re-check, because redesigns and security updates quietly undo this work.