- itscybernews
- Posts
- Cloudflare open-sourced its AI bug hunter. Its best trick: the AI never marks its own homework.
Cloudflare open-sourced its AI bug hunter. Its best trick: the AI never marks its own homework.
One AI hunts for bugs, a second AI tries to prove it wrong, and Cloudflare published the numbers.
Picture this. You hire a hundred detectives to comb through your house for anything a burglar could use. Then you hire a second hundred whose only job is to prove the first hundred wrong. Anything that survives the cross-examination goes on the whiteboard. Everything else gets thrown out.
That is, more or less, how a new open-source tool from Cloudflare works, except the “house” is your source code and the detectives are AI agents. It has been climbing GitHub’s trending lists this week, and it is one of the more interesting ideas to come out of the AI-for-security boom because it is built around a very human problem: nobody trusts a bug report written by a robot that wants to please you.
Meet the skill that argues with itself
Cloudflare’s repository is called security-audit-skill. It is MIT-licensed, and it is a “skill” for a coding agent, which means you install it and then simply tell your agent, in plain English, to “security audit this codebase.” Per the project’s page, installation is a single command from the Skills CLI, and it needs a coding agent that can use tools plus a Node.js runtime.
What happens next is a six-step routine:
Reconnaissance. An agent maps the code’s architecture, its trust boundaries, and every place where outside input comes in.
Coverage-led hunting. Separate agents are each handed one class of vulnerability from a checklist and told to go hunting.
Candidate validation. “Fresh verifiers” with a clean slate try to disprove each suspected bug.
Structured output. Results land in a machine-readable findings.json file, each one marked confirmed, needs_validation, or rejected.
Independent record check. Different agents verify the claims in the final records against the source.
Reporting. Human-readable markdown reports are generated from the verified findings only.
The clever bit is step 3. An AI that finds a bug and then checks its own work tends to agree with itself. So the finder is never the checker. The verifier’s whole job is to say “prove it.” Multiple runs against the same repo are additive, because a coverage ledger remembers what has already been checked.
Does it actually work? Cloudflare showed its homework
The repo grew out of the internal system Cloudflare described in a June blog post, “Build your own vulnerability harness.” The company was unusually candid about the numbers.
What Cloudflare reported | Figure |
|---|---|
Raw candidate bugs the harness produced | 20,799 |
Candidates that survived validation | about 12,057 |
Validation rejection rate, first run vs. later | 40% down to 11% |
Findings folded away as duplicates | 5,442 |
Actionable findings for engineering teams | 7,245 |
Two details are the real story. First, every confirmed finding comes with a proof-of-concept written as a test, run against the original untouched code. A bug that cannot demonstrate itself does not make the list. Second, Cloudflare states plainly that it does not claim a false-negative rate. There is no labelled set of every real bug in a codebase, so any figure for what the tool misses “is entirely speculative.” That is a refreshingly honest sentence in a field full of “99% detection” slide decks.
For scale, the post says a standard repository yields about 100 initial findings per 30,000 lines of code in three to four hours, and a full scan of a complex repository can run past 14 hours.
One quick word from today’s sponsor
Some teams never seem to stop moving. They're on Attio, the agentic CRM.
It’s your always-on revenue engine: agents and workflows build pipeline, chase every buying signal, and move deals forward alongside your team.
Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?
Now the part where it can go wrong
A bug-hunting robot that reads your whole codebase and runs pieces of it is powerful. That is exactly why it deserves a seatbelt.
It executes code. The hunters compile and run fragments of the target code to test their theories. Cloudflare’s own write-up says this belongs in a sandbox, and notes that if you run the harness inside Docker, that container needs seccomp set to unconfined, which loosens one of Docker’s safety features. Run it somewhere disposable, not on your work laptop with your cloud keys sitting in the shell.
Your code goes to a model provider. The skill has no hosted component and needs no Cloudflare service, but your agent still sends source code to whichever AI provider powers it. If your company has rules about that, they apply.
A finding is not a fix. The output is an evidence trail, not a guarantee. Items marked needs_validation are open questions, not conclusions, and the authors say humans still have to apply the threat model and the compliance rules. In Cloudflare’s own automated-patching setup, they call human review of the branch before merge “non-negotiable.”
It costs real money and time. Many isolated agent runs means many tokens. The first pass on a big codebase is the expensive one; later runs get cheaper because the ledger points them at the gaps.
How to try it without regretting it
Start small. Point it at a side project or a public repo you own, not your crown-jewel service.
Use a throwaway box. A fresh VM or container with no secrets, no saved credentials, and network access limited to what the agent needs.
Scrub first. Keep .env files, keys, and customer data out of the folder you audit.
Read the proof. Open the test that comes with each confirmed finding and run it yourself before you believe it.
Treat “needs validation” as a to-do list. Nothing on that list is confirmed either way.
Keep a human on the merge button. Let the machine argue, then let a person decide.
So, is it brilliant or risky?
Both, and that is the charm. The idea at the centre of this tool, that one AI should never mark its own homework, is something a lot of us could use well beyond security. If your bug report has to survive a hostile second opinion before it reaches you, you spend your afternoon on real problems instead of chasing ghosts. Just remember that the same discipline applies to you: run the detectives in a locked room, and check their evidence before you act.
