Cloudflare used frontier AI models inside a controlled harness to probe its web application firewall, turning attacks the firewall had already blocked into new variations. Across 45 scenarios, the system generated 1,107 attempts. Human reviewers retained 49 findings, and the exercise led to three changes in Cloudflare’s Managed Ruleset.
The models operated as black-box testers: they could see responses but had no access to firewall rules, source code or internal security signals. A Python harness—not the models—constructed and replayed HTTP requests, maintained state and enforced limits. One model suggested a mutation while another reviewed the response, allowing later attempts to adapt.
In one server-side request forgery test, the system varied the encoding and placement of a cloud metadata address until it found a representation worth investigation. The setup shows a practical boundary for offensive AI testing: models can explore many variations, but deterministic software controls execution and humans decide which results matter. Cloudflare’s numbers also show the amount of filtering required; most generated attempts did not become validated findings, and only three prompted production rule changes.