STANDARD · §2.1 · DISCOVER · LAST REVIEWED SEPTEMBER 2026

robots.txt for AI crawlers

The oldest file in the stack, and still the first thing to check.


Agent Readiness Compare editors · Spec status checked September 2026

Answer

robots.txt tells crawlers which paths they may fetch, and AI crawlers read it under their own user agent names, such as GPTBot, ClaudeBot and PerplexityBot. It is a request, not access control: RFC 9309 states the rules "are not a form of access authorization." Check that you are not blocking the agents you want by accident, and enforce anything that matters at your server or CDN.

Maintainer:
IETF
Status:
RFC 9309, Standards Track, September 2022

1.What is robots.txt?

robots.txt is a plain text file at the root of a host (/robots.txt) that lists rules for automated clients. Each group starts with one or more User-agent lines followed by Allow and Disallow rules for URL paths. The Robots Exclusion Protocol dates from 1994 and was published as RFC 9309 in September 2022 by M. Koster, G. Illyes, H. Zeller and L. Sassman.

2.How does it apply to AI crawlers?

AI companies publish user agent names for their crawlers and agents, and those names can be targeted in their own groups. A site can allow search crawlers, allow an AI assistant that fetches pages on a user's request, and still disallow training crawlers, by writing separate groups. Rules are matched per user agent, so a broad User-agent: * block also applies to AI crawlers that have no group of their own.

robots.txt (example: allow AI crawlers, keep /account private)
User-agent: *
Disallow: /account/

User-agent: GPTBot
Allow: /
Disallow: /account/

User-agent: ClaudeBot
Allow: /
Disallow: /account/

User-agent: PerplexityBot
Allow: /
Disallow: /account/

Sitemap: https://example.com/sitemap.xml

User agent names change. Confirm current names on each operator's own documentation before relying on them.

3.What does robots.txt not do?

  • It does not enforce anything. RFC 9309: 'These rules are not a form of access authorization.'
  • It does not verify who a crawler is. Any client can send any user agent string. Verification needs signatures (Web Bot Auth) or operator IP lists.
  • It does not summarise your site or expose actions. That is llms.txt, MCP and WebMCP.

4.How do you check it?

  1. Fetch https://yourdomain/robots.txt and confirm it returns HTTP 200 with text content.
  2. Search it for Disallow: / under User-agent: * and under any AI user agent.
  3. Check your CDN or firewall bot rules too. A CDN rule can block a crawler that robots.txt allows. Cloudflare AI Crawl Control can 'track the health of robots.txt files and identify which crawlers are violating your directives'.
  4. Run a readiness scanner; both ora.ai Scan and Cloudflare Is It Agent Ready check robots.txt.

5.Where does it sit in the readiness stack?

Layer 1, Discover: nothing further down the stack works if the agent is turned away here.

Figure 1. The readiness stack: five layers, the standards that address each, and the tools that document support.
Text version of the diagram
Text version of the readiness stack
LayerQuestionStandardsTools that document support
L1 DISCOVERCan an agent find you, and is it allowed in?robots.txt (RFC 9309); Sitemaps; llms.txt; A2A Agent CardCloudflare AI Crawl Control (controls crawler access); ora.ai Scan (checks it); Cloudflare Is It Agent Ready (checks it)
L2 READCan it read and understand what you offer?llms.txt; Markdown negotiation (Accept: text/markdown); JSON-LDCloudflare Markdown for Agents (serves markdown); ora.ai Scan (checks it)
L3 ACTCan it complete a task on your site or API?MCP; WebMCP; agents.json; OpenAPICloudflare (hosts MCP servers); Vercel (hosts MCP servers); nekuda (builds WebMCP tools); ora.ai Journey and WebMCP audit (test it)
L4 PAYCan it pay you?x402; ACP; UCP; MPPCloudflare Pay per crawl (private beta, for crawlers); ora.ai Scan payments layer (checks it)
L5 TRUSTCan you tell which agent it is, and on whose behalf it acts?Web Bot Auth; OAuth 2.0Cloudflare (verifies signed bots); Vercel (verifies signed bots); Forter (links agentic shoppers to verified customer identities)

Frequently asked questions

Should I block AI crawlers in robots.txt?

That is a policy choice. Blocking training crawlers and allowing assistants that fetch pages for a user are different decisions, and robots.txt lets you make them separately by user agent. Blocking everything also blocks agents acting for your prospective customers.

Does robots.txt stop an AI agent from reading my site?

Only if the agent chooses to honour it. RFC 9309 rules are not access authorization. Enforce at the server or CDN.

Source: RFC 9309 (rfc-editor.org) · Cloudflare AI Crawl Control docs · Reviewed Sep 2026