DISPATCH · 18 AUG 2026

How to run a proof of concept for an agent readiness scanner

A two-week plan: agree the criteria, run the tools on the same pages, and compare findings rather than grades.


Agent Readiness Compare editors · Spec status checked September 2026

Agent Readiness Compare editors · · 6 min read

Answer

Agree what success means before you run anything, pick three pages and one task, and run each shortlisted scanner on the same targets in the same week. Compare the findings, the fixes they point to and how well each tool fits your workflow, not the headline grades. Record dates and methods so the write-up survives the next change in the rules.

On this page
  1. 1. Why run a proof of concept at all?
  2. 2. What should you agree before day one?
  3. 3. Which pages and task should you use?
  4. 4. What happens in week one?
  5. 5. What happens in week two?
  6. 6. How do you compare the results?
  7. 7. What are the common traps?
  8. 8. What goes in the write-up?
  9. 9. Sources

1.Why run a proof of concept at all?

Scanner grades are not comparable across tools: each has its own checks, weights and bands. Our guide to choosing a scanner lists the questions to ask on paper; a proof of concept answers them on your own site. It also shows your team what the output looks like before anyone reports a number to leadership.

2.What should you agree before day one?

  • The decision: what the scanner is for, such as finding gaps, gating releases in CI, or reporting to leadership.
  • The targets: three public pages and one task (next section).
  • The criteria and their weights. Start from our six (standards coverage, real-agent task testing, published methodology, fix path, availability and developer workflow) and adjust them in the scanner calculator.
  • The owner and the end date.

3.Which pages and task should you use?

Choose pages an agent needs in order to act for a buyer: the pricing page, one docs or feature page that answers a common question, and the page with your next step, such as a demo request or trial sign-up. For the task, pick one that matters commercially and needs no account, for example: explain our pricing and recommend a plan for a ten-person team. Write the expected answer down first, as lesson 7 describes.

4.What happens in week one?

  1. Day 1: run each scanner on the three pages and save the full reports with the date.
  2. Day 2: run the task. If a tool records agent runs, as ora.ai Journey does, use it; otherwise run the task by hand with a browsing assistant and log each step.
  3. Day 3: put every finding in one table: page, finding, which tools reported it, and which layer it belongs to.
  4. Days 4 and 5: check each finding by hand with the checklist, marking it confirmed or not reproduced.

5.What happens in week two?

  1. Fix two or three confirmed findings, starting with the weakest layer.
  2. Re-run each scanner and the task, and note which tools picked up the fixes.
  3. If CI matters, try each tool's code path. ora.ai documents npx ax audit, a REST API and an MCP server; we found no documented CLI or API for Cloudflare's scanner on its public page.
  4. If you run on Cloudflare, check whether its edge features (Markdown for Agents, AI Crawl Control, Web Bot Auth verification) close findings without code changes. That is where Cloudflare's fix path is strongest.

6.How do you compare the results?

Score each tool against the criteria you agreed, using what you saw rather than what the vendor says. Useful measures: how many confirmed findings each tool reported, how many it missed, how many did not reproduce, whether it showed where the task failed, and how long a re-scan took. Compare the headline grades last, as a sanity check. A tool that reports fewer, better findings can be more useful than one with a longer list.

7.What are the common traps?

  • Comparing grades across tools. They use different scales.
  • Testing only the home page. Agents need pricing, docs and forms.
  • Forgetting login-gated flows. Scanners of this kind read public URLs, and ora.ai states that it cannot evaluate login-gated content, so test those flows separately.
  • Changing the site mid-test without noting it. Date every change.
  • Treating a score as an outcome. No scanner measures sign-ups or revenue.

8.What goes in the write-up?

One page: the decision, the targets, the criteria and weights, a table of findings by tool, the task results, what was fixed and whether each tool detected the fix, and a recommendation. Attach the dated reports. If the choice is close, say so. Many teams end up with one tester and one edge provider, because the two do different jobs; see testing tools and fixing tools.

9.Sources

Reviewed Sep 2026