AI NewsInfrastructureAnnouncement
Cloudflare ran 1,107 attack attempts against its own WAF with an AI tester and kept 49 findings after review
Cloudflare published on 29 September that it built an AI-driven WAF tester which ran 1,107 attempts across six attack categories on an authorised staging environment, and left 49 findings worth investigating after human review, 48 of them in command injection and server-side request forgery.

Image: Cloudflare
Why it mattersA security team can copy this pattern to test its own WAF configuration against attack payloads a static test list would miss, but the real number to plan around is 49 findings from 1,107 attempts, which is how much human review the model still needs.
A security team that wants to test whether its web application firewall stops a frontier AI model producing attack payloads now has a public description of what the numbers look like. Cloudflare published a post on 29 September saying its own team ran an AI-driven tester against a customer staging environment, recorded 1,107 attempts, and kept 49 findings worth investigating after human review, 48 of which fell in command injection and server-side request forgery.
The tester starts from an attack the WAF already blocks, sends variations of it, and uses each response to pick the next variation. Cloudflare says two model calls run each loop: one proposes the next request from a short history of earlier results, and the second reviews the response status, selected headers and body. The loop stops when the model runs out of new variations or hits an attempt limit.
What the run covered
Cloudflare ran 45 scenarios. 44 covered six categories: cross-site scripting, SQL injection, command injection, server-side request forgery, path traversal or local file inclusion, and Log4j. The remaining scenario covered log injection and is reported separately. The WAF in the test zone was configured with WAF Attack Score blocking at scores of 30 or below, the full Cloudflare Managed Ruleset enabled, and the OWASP Core Ruleset at Paranoia Level 3.
The tester ran on an authorised customer staging environment, with a test User-Agent on the allowlist so automated-traffic controls would not stop the run. Cloudflare says the models never received rule expressions, rule IDs, or WAF Attack Score details, and that the code, not the model, controlled request replay and hostname allowlisting.
What the numbers said
XSS, LFI, SQL injection and Log4j got near full coverage. Command injection and server-side request forgery held 48 of the 49 findings kept after review. Cloudflare also filtered out attempts where the model failed to produce a valid HTTP request, the request never reached the target, the payload had lost its malicious form during mutation, or the behavior did not belong to the WAF at all.
Cloudflare gives one concrete example. In an SSRF scenario, the tester sent the same cloud metadata address in different forms, including integer, octal, and a trailing-dot version of the IP, and placed it in different parts of the request. The WAF blocked all but one. On attempt 18, the model kept the previous request structure and switched to the trailing-dot form, and the client got a redirect rather than a WAF block. That finding shipped as a new SSRF Obfuscated Host detection in Cloudflare's July 21 Managed Ruleset release, together with an SSRF Restricted Protocol detection and an improvement to the existing SSRF Cloud rule.
What Cloudflare says the run did not prove
Two things stand out in the post. Cloudflare ran the same scenarios against two versions of one model family and got different variations, but the underlying issues appeared in both, so it treated the models as attack generators and not as a source of truth. And extending an attempt limit inside one scenario returned less than testing more starting requests, more categories, and more input positions.
Cloudflare notes that this was black-box testing: the models saw no source code, no WAF rules, and only selected response data. A follow-up post is planned covering the white-box case, where the model sees both the application's vulnerabilities and the WAF rules protecting it.
Source
- Primary source: We tested our own WAF with frontier AI models. Here's what we found, Cloudflare blog, 29 September 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


