For twenty years, websites were built for two audiences: people, and search engine crawlers. A third has arrived. AI agents such as xAI's Grok Bot and OpenAI's ChatGPT Agent open your site in a real browser and do things on it for their users: compare products, fill in quote forms, check availability, book appointments.
When an agent fails on your site, your competitor gets the order, and you never find out why. This guide shows what agents need, what stops them, and how to fix it, in order of impact. Every check below is also in the free AI Agent Readiness Check, so you can test as you go.
Key Takeaways
- AI agents use a real browser from a cloud server. They ignore robots.txt and fail on the same things as an impatient human.
- The biggest blockers are firewall challenges on data-centre traffic and CAPTCHAs on everyday forms.
- Accessible HTML is agent-friendly HTML: real buttons, labelled fields, real links.
- Keep key facts in the HTML, not only behind JavaScript.
- Keep human checks on payment and account creation. That is where they belong.
How AI agents use a website
An AI agent is not a crawler. A crawler collects pages in bulk for an index. An agent opens a few specific pages to complete one task for one person, then leaves.
xAI's Grok Bot overview describes the model plainly: each Bot "works on a persistent cloud computer with a browser, filesystem, and terminal", using connectors where they exist "and computer use for everything else". Computer use means the agent looks at the page, decides what to click or type, and does it, the way a person would.
Three facts follow from that, and they shape everything else in this guide.
- Agents come from data centres. They run on cloud computers, so your firewall sees cloud-provider IP addresses, not home broadband.
- Agents read the page structure. They work out what a button does from its label and markup. Ambiguous or fake controls slow them down or mislead them.
- Agents hand human steps back. xAI's Grok Bot FAQ says the Bot should hand CAPTCHAs, logins and confirmations back to its user. Every such step is a pause, and some users will not come back.
If your site is accessible and fast for a person on a slow connection, it is most of the way to being good for agents. The rest is about not mistaking them for attackers.
Step 1: Let cloud browsers reach your pages
This is the biggest single blocker. Many sites, often without knowing it, challenge or refuse traffic from cloud providers, because that is where most scraping comes from. xAI's Grok Bot security FAQ names it as the reason "some websites block Bots": "Some services flag datacenter IP addresses."
- Check your bot-management level. On Cloudflare, Akamai and similar services, "high" or "under attack" modes challenge almost all data-centre traffic. Use them on login and checkout only.
- Look for blanket ASN or country blocks. A rule blocking a whole cloud provider stops every agent hosted there.
- Prefer rate limits to blocks. A limit on requests per minute stops scraping without refusing a single agent that loads five pages.
The readiness check loads your page from a cloud server, once as a Linux browser and once as a desktop browser, and flags both refusals and a different answer for the two. Our guide to firewalls that block AI by accident covers the most common rules.
Step 2: Move CAPTCHAs to where they belong
CAPTCHAs exist to stop automated abuse, so it is no surprise that they stop automated assistance too. The fix is not to remove them, but to stop using them where no abuse is possible.
| Page or action | CAPTCHA? | Why |
|---|---|---|
| Search, browsing, product pages | No | Agents research here. Rate limit instead. |
| Contact and quote forms | Only risk-based or invisible | These are leads. Filter spam after submission. |
| Newsletter sign-up | Optional | Use double opt-in, which stops abuse without a puzzle. |
| Account creation | Yes | Prevents fake accounts. |
| Payment and checkout | Yes, if you already use one | A human is present to approve anyway. |
| Password reset | Yes | A security action. |
Risk-based modes, such as invisible reCAPTCHA, Cloudflare Turnstile in managed mode or hCaptcha's passive mode, only show a challenge to sessions that look suspicious. That keeps most legitimate agents moving. The readiness check reports any of these widgets it finds in the HTML; one injected later by script will not appear, so test forms by hand too.
Step 3: Put the facts in the HTML
Agents run JavaScript, unlike most AI crawlers. But a page that shows nothing until several scripts have loaded is slower and less reliable for them, and invisible to the crawlers that feed AI answers in the first place.
- Render prices, stock, opening hours, delivery times and contact details on the server.
- Avoid "click to reveal" for information people ask about most.
- Keep the main content in the first HTML response, then enhance it with scripts.
A quick test: fetch the page with JavaScript off, or with curl, and see whether the facts are there. Our guide on JavaScript rendering and AI crawlers explains the trade-offs for different frameworks.
Step 4: Use real buttons, links and labels
An agent decides what to do from what the page tells it. The same markup that makes a site work with a screen reader makes it work for an agent, because both rely on the accessibility tree rather than the visual design.
Buttons and links
- Use
<button>for actions and<a href>with a real URL for navigation. - Avoid clickable
<div>and<span>elements, and links to#orjavascript:. - Give icon-only buttons an
aria-label, such as "Add to basket".
This is the same requirement as WCAG 2.2 "Name, Role, Value": every control needs a name and a role that software can read.
Form fields
<label for="qty">Quantity</label>
<input id="qty" name="quantity" type="number" min="1" value="1" autocomplete="off">
- Every field has a visible
<label for>. Placeholders alone disappear once typing starts, which confuses agents and people alike. See WCAG 2.2 "Labels or Instructions". - Use the right
type(email,tel,date,number) andautocompletevalues, so the agent knows what each field wants. - Show errors next to the field, in text, not only as a red border.
The readiness check counts the fields on your page and reports how many have a label, an aria-label or a placeholder, and how many clickable elements are not real buttons or links.
Step 5: Make important flows short and predictable
Agents are good at following a clear path and bad at guessing. Every extra modal, surprise redirect and layout change is a place to get lost.
- Fewer steps. A quote form on one page beats a five-step wizard.
- Stable URLs. Let a product, a filtered search or a booking date live at its own URL, so an agent can go straight to it.
- No interstitials on arrival. Cookie banners are fine; a full-screen newsletter pop-up on every page is a wall.
- Clear confirmation. After a submission, show what happened in text on the page, and send a confirmation email. The agent can report success, and the human behind it can check.
Step 6: Keep humans in the loop where it matters
Agent-friendly does not mean agent-only. Some actions should always need a person, and xAI's design agrees: its approvals page tells users to type passwords, two-factor codes and CAPTCHA answers themselves, and its security documentation advises keeping sending, purchasing and publishing behind approval.
So keep your human checks on payment, account creation, password changes and anything irreversible. Make them clear and quick, because a person will be completing them, often from a phone, in the middle of a task their agent started.
Step 7: Offer a better path than clicking
xAI says Grok Bot uses connectors "where available" and computer use for everything else. A connector or API is faster and more reliable for the agent, and puts you in control of exactly what it can do.
- If you already have a public API, document it clearly and link to it from your site.
- If your customers use AI tools heavily, consider an MCP server. Our guide to MCP and APIs for AI agents explains the options.
- Add structured data such as schema.org Product to product pages, so prices and availability can be read without parsing the layout.
What changes for different kinds of site
The steps above apply everywhere, but the pages that matter most differ by business. Start with the ones in your row.
| Type of site | Pages agents use most | Most common blocker |
|---|---|---|
| Online shop | Search, category, product, basket | CAPTCHA on search; prices loaded by script |
| Hotel, restaurant, venue | Availability, booking, menus | Booking engine on a third-party domain behind its own firewall |
| B2B supplier or manufacturer | Product specs, quote form, distributor finder | Quote form CAPTCHA; specs only in PDFs |
| Local service business | Services, areas covered, contact or booking | Strict firewall on a small site; contact details in an image |
| SaaS product | Pricing, docs, sign-up, in-app pages | Pricing behind "contact sales"; docs rendered only by script |
| Publisher | Articles, archives, account | Full-page pop-ups; paywall that hides even the headline |
One pattern deserves its own warning: third-party booking and checkout systems. If your booking widget runs on another company's domain, its firewall settings apply, not yours. Test the full journey, including the part that leaves your site, and ask your provider how they treat automated browsers.
Measuring agent traffic
You cannot improve what you cannot see, and agent visits are easy to miss because they look like ordinary browser sessions. Three ways to get a rough picture:
- Network or ASN reports. Most analytics tools can break visits down by network. A rise in sessions from cloud providers, with short, direct paths to forms, is a good sign of agent use.
- Firewall events. Your CDN's security log shows challenges and blocks. Many challenges on public pages from cloud networks means you are turning agents away.
- Form completion by network. If forms started from cloud networks are rarely completed, the form itself is the likely blocker.
For the crawler side, which is easier to measure because crawlers send their names, see tracking AI bots in server logs and, for xAI names specifically, the Grok log checker.
Agents, crawlers and AI answers: how they fit together
It is easy to treat "AI" as one thing. For a website it is at least three, and each needs its own setting.
| What | Example | What it needs from you |
|---|---|---|
| Training crawlers | GPTBot, ClaudeBot | A robots.txt decision: allow or opt out |
| Search and answer crawlers | OAI-SearchBot, xAI-SearchBot | Allowed in robots.txt and readable without JavaScript |
| Browsing agents | Grok Bot, ChatGPT Agent | Reachable from the cloud, no needless human checks, clear markup |
A site can block training, allow search and welcome agents, and that is a perfectly consistent policy. Our guide to blocking AI training while allowing AI search covers the crawler half; this guide covers the agent half.
A one-hour audit you can run today
If you only have an hour, this order gives the most value for the time.
- Ten minutes: list the five pages where customers act: search, a top product or service page, booking or basket, contact or quote, and pricing.
- Fifteen minutes: run each through the AI Agent Readiness Check and note every fail.
- Fifteen minutes: open your CDN or firewall settings and look at the security level and any rules that mention bots, countries or cloud networks. Note which apply to the five pages.
- Ten minutes: complete each form yourself on a phone, with a VPN on. Where you hit a challenge, an agent will too.
- Ten minutes: write the fix list, starting with anything that refuses or challenges the page outright.
Most sites find one or two hard stops in that hour. Fixing those usually matters more than everything else in this guide put together.
Write down what you changed and when. Agent behaviour and firewall vendors both change often, and a dated record makes it easy to tell whether a later drop in form completions came from your change or from theirs. Repeat the audit after every redesign, every new form, and every change of CDN or security provider, because those are the moments when a working page most often stops working for agents without anyone noticing.
What not to do
- Do not rely on robots.txt. Agents act for people and do not read it. robots.txt still matters for crawlers such as GrokBot; see the Grok crawler guide.
- Do not hide content from agents. Serving different content to automated browsers is fragile, and it breaks for real users on VPNs and corporate networks too.
- Do not put instructions to agents in your pages. xAI treats web content as untrusted data, and hidden text aimed at AI can look like manipulation to people and to search engines.
- Do not block whole cloud providers. It stops every agent and many legitimate services, and is rarely what you meant.
Questions teams ask before changing anything
Will lowering our firewall settings let in more attacks?
Not if you lower them selectively. Keep strict settings on login, admin, checkout and APIs, where attacks happen. Public marketing and product pages carry little risk, and rate limits protect them better than challenges do, because a limit slows a scraper without stopping a single shopper or agent.
Our CAPTCHA blocks a lot of spam. What replaces it?
Spam filtering after submission, honeypot fields that only bots fill in, and rate limits per IP or session. These stop most form spam without asking anyone to solve a puzzle. Keep a challenge in reserve for sessions that trip those filters.
How do we know whether agents matter for us yet?
Look at your firewall and analytics for cloud-network traffic on forms, as described above, and ask your sales or bookings team whether customers mention AI assistants. Agent use is growing fastest in research-heavy purchases, such as B2B supplies, travel and services, where people delegate comparison work.
Does any of this affect SEO?
Indirectly, and only for the better. Labelled forms, real links, server-rendered facts and fewer pop-ups are all things search engines and accessibility audits already reward. Nothing in this guide asks you to change robots.txt or hide content from crawlers.
How the readiness check scores your page
The AI Agent Readiness Check runs eight checks on one page and scores it from 0 to 100. A pass counts fully, a "check" counts half, a fail counts zero, and information-only results are not scored. The worst problem then caps the score, so one serious issue cannot hide behind easy passes: a refusal or firewall challenge caps it at 30, any other fail at 60, a CAPTCHA or login wall at 75, and any other warning at 85.
| Check | What passes |
|---|---|
| Reachable from a cloud browser | HTTP 200 from a data-centre server |
| No firewall challenge | No challenge or block page from the major WAF vendors |
| No CAPTCHA on the page | No CAPTCHA widget in the HTML |
| Content visible without logging in | Public content, not just a password form |
| Content in the HTML | At least 800 characters of readable text without JavaScript |
| Labelled form fields | 90% or more fields with a label, aria-label or placeholder |
| Real buttons and links | No clickable divs or dead links |
| Same response for automated browsers | Same HTTP status for a Linux and a desktop browser |
It is a strong signal, not proof. The check runs from our servers, not from any agent's network, and reads raw HTML without running scripts or submitting forms. A site that recognises real agents by IP or signature can treat them differently. Run it on the pages that matter most: search, product, booking, contact and quote.
A worked example
Here is an illustrative example of how a small business might work through this guide. A regional plumbing company gets most of its jobs from a "Request a quote" form. A customer asks their AI agent to "get quotes from three plumbers near me for a boiler service next week".
- The agent finds the company's quote page. The firewall is set to "high", so the agent gets a challenge and hands it back to the customer, who is busy and moves on. Fix: lower the security level for the quote page and keep "high" on the admin area.
- The form has a reCAPTCHA checkbox. Fix: switch to the invisible version and filter spam after submission.
- The postcode and date fields have placeholders but no labels, and the "Next" button is a styled div. Fix: add labels and make it a real button. This also fixes two accessibility issues.
- After submission the page shows a spinner and nothing else. Fix: show a confirmation message with the reference number and send an email.
Each of these is an afternoon's work or less, and every one also helps human visitors. Fixes 1 and 2 remove the two hard stops; fixes 3 and 4 remove the friction that makes an agent slow or uncertain.
Your agent-readiness checklist
- □Run the AI Agent Readiness Check on your five most important pages.
- □Lower firewall challenge levels on public pages, and keep them on login, admin and checkout.
- □Move CAPTCHAs off search and enquiry forms, or switch them to risk-based modes.
- □Render prices, availability and contact details in the HTML.
- □Replace clickable divs with real buttons and links, and label every form field.
- □Show a clear text confirmation after every form submission.
- □Keep human checks on payment, account creation and password changes.
- □Re-test after every redesign, because agents rely on structure that redesigns often change.
Agent readiness is not a one-off project. It is the same discipline as accessibility and page speed: small, steady habits that keep a site usable as it changes. The difference is that the visitor you are designing for now includes software working for your customers.
For the blocking side of the same question, read can you block Grok Bot?. For the difference between agents and crawlers, read Grok Bot vs GrokBot.
