Industry case study

How Cloudflare made AI security findings easier to verify

Cloudflare paired AI investigation with a structured review process. The lesson is how a team moves from a suspected flaw to evidence it can act on.

Source publishersAnthropic / Cloudflare
Source published22 May 2026
Last checked

Independent Cactera analysis of publicly documented work. Cactera was not involved in this work. Company names identify the subjects, not Cactera clients or partners.

Company result reported by Anthropic

2,000 bugsreported across Cloudflare’s critical-path systems

Read Anthropic’s account
Comparison
No independent, like-for-like comparison published
Scope
Cloudflare systems; 400 of the reported bugs rated high or critical
Timeframe
First-month Project Glasswing update, 22 May 2026
The published work

The problem.

Cloudflare found that asking a general coding agent to inspect a repository produced noisy findings and limited useful coverage. Reviewers still needed to understand which issues mattered.

What changed.

Its team tested Mythos Preview across more than 50 repositories, adding focused investigations, separate validation, duplicate removal and checks on whether affected code could be reached.

As described by Anthropic and Cloudflare.

Company result reported by Anthropic

What was reported.

Anthropic’s May update reports 2,000 bugs found across Cloudflare’s critical-path systems, including 400 rated high or critical. Cloudflare’s own article explains the research process behind the work. Source: Anthropic

What the evidence can tell us

The count is reported by the model provider; it is not a count of prevented attacks or proof that every bug was exploitable or patched. Cloudflare reports remaining false positives and some AI-written fixes that introduced regressions.

Cactera analysis

What we take from it.

For a cloud application, a useful review starts with what the deployed system actually allows. We would identify the services, their owners and the expected boundaries between them before choosing a tool. This gives a reviewer a way to judge whether an apparent weakness changes what a real user or service can do.

A finding also needs a decision path. One engineer may understand the affected library while another owns the application that uses it. We would make that handoff explicit, with a shared issue that carries the relevant evidence and names the next person responsible. Otherwise, technically credible findings can wait indefinitely between teams.

The queue should separate uncertainty from urgency. An incomplete investigation needs a next step, rather than a confident severity label that obscures what is missing. We would track time spent confirming, dismissing and fixing issues as well as the findings themselves. That makes it possible to see whether AI reduces review effort or simply moves it elsewhere.

The final check belongs in the application’s normal operating context. A useful pilot should end with reviewed changes, observable deployment and an owner for unresolved work. We would compare the effort against the same team’s existing process before expanding it. The objective is a maintainable route from evidence to resolution that continues to work after the initial assessment.

A proposed method for your business

How to evaluate a similar idea.

Start with your situation and a question you can test. These are evaluation steps we would discuss before choosing an implementation.

  1. 01

    Map the deployed system

    Identify service ownership, expected access and the operational context a reviewer needs to assess a finding.

  2. 02

    Make handoffs explicit

    Give each issue an owner and a next decision, particularly when a shared dependency affects several teams.

  3. 03

    Measure review effort

    Track time spent on confirmation, dismissal and remediation so the evaluation includes the work generated by the tool.

  4. 04

    Close the operational loop

    Review fixes, check normal behavior and confirm deployment before recording an issue as resolved.

Industry case study / Source notes

Sources & credits.

Work credited to
Cloudflare’s security and engineering teams
Technology / platform
Anthropic
Analysis & explanation
Cactera. Company wordmarks identify the article subjects.

Independent Cactera analysis of publicly documented work. Cactera was not involved in this work. Company names identify the subjects, not Cactera clients or partners.

A relevant next step

Bring the right question.
Let’s make it specific.

Explore how cloud security could fit the work you have in mind.

Get a quote