AI NewsInfrastructureAnnouncement
Cloudflare used an AI harness to find three WAF gaps
Cloudflare let frontier AI models rewrite blocked attacks on a staging site, ran 1,107 attempts across 45 scenarios, and shipped three Managed Ruleset changes from the surviving 49 findings.

Image: Cloudflare
Why it mattersA security team now has a reproducible pattern for using AI to probe a defence it already runs, with the model constrained to proposing payload mutations and the harness owning everything that actually hits the service.
A fixed set of security tests runs the same payloads every time, and the gaps a WAF has against variations of those payloads stay closed. Cloudflare's security team wrote up what it found when it placed frontier AI models inside a controlled harness and let them propose variations instead. The post went up on 29 September.
The run produced three changes to Cloudflare's Managed Ruleset and a reusable design that a team can run against its own service.
Where control stayed
A Python harness built and replayed the HTTP requests, maintained scenario state, enforced limits, and collected the responses. Two model calls per round: one proposed the next mutation to the payload, and a second reviewed the response from the previous attempt. The models never saw the WAF rules, source code, or internal security signals, so the test was black-box from their point of view.
The run started from attack payloads the WAF already blocked. On each round, the model proposed a change: a different encoding, a different placement in the request, or the same destination written another way. Cloudflare ran the harness against an authorised customer staging environment with its Managed Ruleset enabled and OWASP Core Ruleset at Paranoia Level 3.
The numbers from the run
Cloudflare says the run covered 45 scenarios across six attack categories: XSS, SQL injection, command injection, server-side request forgery, path traversal and local file inclusion, and Log4j. It recorded 1,107 attempts in total. After the team removed malformed, benign, duplicate, and out-of-scope results, 49 findings survived human triage. Forty-eight of those 49 were in command injection or SSRF; XSS, LFI, SQLi, and Log4j had near full coverage already.
A worked SSRF example in the post describes the loop. The tester kept changing how a cloud-metadata address was written: decimal, octal, and other forms. Decimal was blocked. On the next attempt, the model kept the request shape and switched to a trailing-dot representation of the address, and the request went to a redirect rather than a block.
What shipped
Three changes landed in the 21 July Managed Ruleset release: two new detections, SSRF - Obfuscated Host and SSRF - Restricted Protocol, and an improvement to the existing SSRF - Cloud rule. Cloudflare says the SSRF - Obfuscated Host detection came directly from requests that encoded internal addresses in non-standard numeric forms.
A security team that already runs tests against its own service has a template here. The model proposes variations the harness controls, the harness collects evidence under limits the humans set, and the review stays with a reviewer who decides whether a surviving request is actually a bypass or just a strange shape that reached the backend without doing anything.
Source
Primary: We tested our own WAF with frontier AI models. Here's what we found, Cloudflare Blog.
Reporting: Cloudflare Uses an AI Harness to Probe and Harden Its WAF, Matt Foster, InfoQ.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.