UNIVERSITY · LESSON 9 · RUNNING IT · 5 MIN READ

Setting a policy for AI crawlers and agents

Training, search and assistants are three decisions, not one.


Agent Readiness Compare editors · Spec status checked September 2026

Answer

Decide separately whether to allow AI training crawlers, search crawlers and assistants that fetch pages for a user, then express each decision in robots.txt groups and enforce what matters at your CDN. Use verified bot identity rather than user agent names where you can, and decide whether any content should be paid for rather than free or blocked.

On this page
  1. 1. Which decisions does a policy cover?
  2. 2. How do you express the policy?
  3. 3. How do you enforce it?
  4. 4. Should any content be paid for?
  5. 5. How do you keep the policy current?

1.Which decisions does a policy cover?

  • Training: may AI companies use your content to train models?
  • Search: may AI search and answer engines index your pages?
  • Assistants and agents: may an agent acting for a user, such as a prospective customer, fetch and use your pages?

Blocking everything answers all three with no, including the last, which turns away agents working for your buyers. Writing the three decisions down separately is the first step.

2.How do you express the policy?

robots.txt lets you write a group per user agent, so a site can allow search crawlers and assistants and disallow a training crawler by name. RFC 9309 is clear that these rules "are not a form of access authorization": well-behaved crawlers follow them, and others may not. See robots.txt for AI crawlers for an example file.

CDNs add settings on top. On 15 September 2026 Cloudflare introduced a "Disallow AI Training" setting that uses robots.txt Disallow rules to opt out of training while staying discoverable in search, with an "Accountable" designation for crawler operators that meet stated opt-out and reporting requirements. Cloudflare's Markdown for Agents sends a default content-signal: ai-train=yes, search=yes, ai-input=yes header unless the origin overrides it, so check that header against your policy before enabling it.

3.How do you enforce it?

At the edge, with verified identity. User agent strings can be copied, so a rule that trusts a name also trusts anything that claims it. Web Bot Auth lets a CDN check a signature instead: Cloudflare documents it as a bot verification method, and Vercel's bot verification supports it. Cloudflare AI Crawl Control, available on all plans, shows which AI services crawl a site, lets you allow or block individual crawlers, and tracks which crawlers violate robots.txt directives.

4.Should any content be paid for?

A third option sits between allow and block: charge. Cloudflare's Pay per crawl, part of AI Crawl Control and in private beta, lets AI crawlers pay to access content. For agents that buy data or API calls per request, x402 uses HTTP 402 so a client can pay and retry. Both are commercial decisions; record them in the same written policy (checklist L4.4).

5.How do you keep the policy current?

Crawler and agent names change, so confirm current names on each operator's own documentation, review the policy whenever you add a CDN rule, and re-test that the agents you want can still reach your key pages, as lesson 8 describes. That completes the University; the checklist turns all of it into 24 checks.

Source: RFC 9309 · Cloudflare blog, Accountable mixed-use AI crawlers · Cloudflare Markdown for Agents docs · Cloudflare Web Bot Auth docs · Vercel verified bots · Cloudflare AI Crawl Control docs · Cloudflare Pay per crawl docs · x402.org · Reviewed Sep 2026

Back to University