STANDARD · §2.1 · DISCOVER · LAST REVIEWED SEPTEMBER 2026
robots.txt for AI crawlers
The oldest file in the stack, and still the first thing to check.
Agent Readiness Compare editors · Spec status checked September 2026
Answer
robots.txt tells crawlers which paths they may fetch, and AI crawlers read it under their own user agent names, such as GPTBot, ClaudeBot and PerplexityBot. It is a request, not access control: RFC 9309 states the rules "are not a form of access authorization." Check that you are not blocking the agents you want by accident, and enforce anything that matters at your server or CDN.
- Maintainer:
- IETF
- Status:
- RFC 9309, Standards Track, September 2022
- Spec:
- rfc-editor.org
1.What is robots.txt?
robots.txt is a plain text file at the root of a host (/robots.txt) that lists rules for automated clients. Each group starts with one or more User-agent lines followed by Allow and Disallow rules for URL paths. The Robots Exclusion Protocol dates from 1994 and was published as RFC 9309 in September 2022 by M. Koster, G. Illyes, H. Zeller and L. Sassman.
2.How does it apply to AI crawlers?
AI companies publish user agent names for their crawlers and agents, and those names can be targeted in their own groups. A site can allow search crawlers, allow an AI assistant that fetches pages on a user's request, and still disallow training crawlers, by writing separate groups. Rules are matched per user agent, so a broad User-agent: * block also applies to AI crawlers that have no group of their own.
User-agent: *
Disallow: /account/
User-agent: GPTBot
Allow: /
Disallow: /account/
User-agent: ClaudeBot
Allow: /
Disallow: /account/
User-agent: PerplexityBot
Allow: /
Disallow: /account/
Sitemap: https://example.com/sitemap.xmlUser agent names change. Confirm current names on each operator's own documentation before relying on them.
3.What does robots.txt not do?
- It does not enforce anything. RFC 9309: 'These rules are not a form of access authorization.'
- It does not verify who a crawler is. Any client can send any user agent string. Verification needs signatures (Web Bot Auth) or operator IP lists.
- It does not summarise your site or expose actions. That is llms.txt, MCP and WebMCP.
4.How do you check it?
- Fetch https://yourdomain/robots.txt and confirm it returns HTTP 200 with text content.
- Search it for
Disallow: /underUser-agent: *and under any AI user agent. - Check your CDN or firewall bot rules too. A CDN rule can block a crawler that robots.txt allows. Cloudflare AI Crawl Control can 'track the health of robots.txt files and identify which crawlers are violating your directives'.
- Run a readiness scanner; both ora.ai Scan and Cloudflare Is It Agent Ready check robots.txt.
5.Where does it sit in the readiness stack?
Layer 1, Discover: nothing further down the stack works if the agent is turned away here.
The readiness stack: five layers (Discover, Read, Act, Pay, Trust), the standards that address each, and the tools that document support.
Can an agent find you, and is it allowed in?
Can it read and understand what you offer?
Can it complete a task on your site or API?
Can it pay you?
Can you tell which agent it is, and on whose behalf it acts?
L1 DISCOVER
Can an agent find you, and is it allowed in?
Cloudflare AI Crawl Control (controls crawler access); ora.ai Scan (checks it); Cloudflare Is It Agent Ready (checks it)
L2 READ
Can it read and understand what you offer?
Cloudflare Markdown for Agents (serves markdown); ora.ai Scan (checks it)
L3 ACT
Can it complete a task on your site or API?
Cloudflare (hosts MCP servers); Vercel (hosts MCP servers); nekuda (builds WebMCP tools); ora.ai Journey and WebMCP audit (test it)
L4 PAY
Can it pay you?
Cloudflare Pay per crawl (private beta, for crawlers); ora.ai Scan payments layer (checks it)
L5 TRUST
Can you tell which agent it is, and on whose behalf it acts?
Cloudflare (verifies signed bots); Vercel (verifies signed bots); Forter (links agentic shoppers to verified customer identities)
| Layer | Question | Standards | Tools that document support |
|---|---|---|---|
| L1 DISCOVER | Can an agent find you, and is it allowed in? | robots.txt (RFC 9309); Sitemaps; llms.txt; A2A Agent Card | Cloudflare AI Crawl Control (controls crawler access); ora.ai Scan (checks it); Cloudflare Is It Agent Ready (checks it) |
| L2 READ | Can it read and understand what you offer? | llms.txt; Markdown negotiation (Accept: text/markdown); JSON-LD | Cloudflare Markdown for Agents (serves markdown); ora.ai Scan (checks it) |
| L3 ACT | Can it complete a task on your site or API? | MCP; WebMCP; agents.json; OpenAPI | Cloudflare (hosts MCP servers); Vercel (hosts MCP servers); nekuda (builds WebMCP tools); ora.ai Journey and WebMCP audit (test it) |
| L4 PAY | Can it pay you? | x402; ACP; UCP; MPP | Cloudflare Pay per crawl (private beta, for crawlers); ora.ai Scan payments layer (checks it) |
| L5 TRUST | Can you tell which agent it is, and on whose behalf it acts? | Web Bot Auth; OAuth 2.0 | Cloudflare (verifies signed bots); Vercel (verifies signed bots); Forter (links agentic shoppers to verified customer identities) |
Text version of the diagram
| Layer | Question | Standards | Tools that document support |
|---|---|---|---|
| L1 DISCOVER | Can an agent find you, and is it allowed in? | robots.txt (RFC 9309); Sitemaps; llms.txt; A2A Agent Card | Cloudflare AI Crawl Control (controls crawler access); ora.ai Scan (checks it); Cloudflare Is It Agent Ready (checks it) |
| L2 READ | Can it read and understand what you offer? | llms.txt; Markdown negotiation (Accept: text/markdown); JSON-LD | Cloudflare Markdown for Agents (serves markdown); ora.ai Scan (checks it) |
| L3 ACT | Can it complete a task on your site or API? | MCP; WebMCP; agents.json; OpenAPI | Cloudflare (hosts MCP servers); Vercel (hosts MCP servers); nekuda (builds WebMCP tools); ora.ai Journey and WebMCP audit (test it) |
| L4 PAY | Can it pay you? | x402; ACP; UCP; MPP | Cloudflare Pay per crawl (private beta, for crawlers); ora.ai Scan payments layer (checks it) |
| L5 TRUST | Can you tell which agent it is, and on whose behalf it acts? | Web Bot Auth; OAuth 2.0 | Cloudflare (verifies signed bots); Vercel (verifies signed bots); Forter (links agentic shoppers to verified customer identities) |
Frequently asked questions
Should I block AI crawlers in robots.txt?
That is a policy choice. Blocking training crawlers and allowing assistants that fetch pages for a user are different decisions, and robots.txt lets you make them separately by user agent. Blocking everything also blocks agents acting for your prospective customers.
Does robots.txt stop an AI agent from reading my site?
Only if the agent chooses to honour it. RFC 9309 rules are not access authorization. Enforce at the server or CDN.
Source: RFC 9309 (rfc-editor.org) · Cloudflare AI Crawl Control docs · Reviewed Sep 2026