The Machines Got to 32 of 36. Humans Finished It.
A competitive benchmark let AI agents run openly against the best human security teams in the world. The agents did well. The team that finished everything was still human, and that gap is now a marketing problem.
By Katie Delaney · 2026-08-29 · 13 min read
Somebody finally ran the scoreboard#
If you sell security services, you have spent a year being told that automated penetration testing is about to make your delivery team redundant. Last week somebody finally ran the scoreboard, and the scoreboard is more interesting than either the hype or the backlash.
Hack The Box ran its 2026 Global Cyber Skills Benchmark under the name Project Nightfall and let agent teams compete openly. Help Net Security's report on the resulting document gives the numbers: agent accounts "held 2.7% of registered accounts and produced 4.2% of submitted flags and 4.6% of awarded points".
The headline result is the one nobody put in a press release#
"The team that finished everything, all 36 challenges, was human. The best AI team stopped at 32." Four challenges is not a rounding error at the top of a competitive field. It is the difference between solving the problem and solving most of the problem, which in this trade is the difference that matters.
The participation figures are worth having too, because they set the scale honestly. There were "93 designated agent accounts across 54 teams, 46 of them active". Thirty-three of the Top 100 teams had an agent, and agents appeared in 17 of the Top 25 finishers. So agents are genuinely competitive near the top. They are also, at present, a garnish rather than the meal.
Two framings are available and only one of them survives contact with a buyer. The first says automated penetration testing has arrived and the humans are finished. The second says the whole thing is hype and nothing has changed. The measured answer is neither: agents outperformed their headcount by roughly half again, and still did not close the course.
Worth noting what kind of test this is. A benchmark is an artificial environment with solvable problems and a scoreboard, which flatters automation. Real automated penetration testing work happens in estates that are undocumented, half-broken and full of systems nobody owns. If agents fall four short under favourable conditions, the honest expectation in a client environment is wider, not narrower.
The trail here is short and worth walking properly. One benchmark, one independent paper, one week of practitioner reaction, and a great deal of marketing built on top of all three. Automated penetration testing is the rare category where the evidence is public and the claims still outrun it.
Three caveats before you quote any of it#
Before anyone builds a campaign on this, three caveats, because the sourcing here is thinner than the confidence around it.
What we could and could not verify#
First, the underlying document. The figures above are quoted from Help Net Security, which attributes them explicitly to Hack The Box's own report. We tried to read that report directly and could not: the file exceeds our fetching limits and would not parse. So these numbers are trade press quoting a vendor report, verified against one another and internally consistent, but not read at source by us. That is worth saying out loud rather than implying otherwise.
Second, a figure that is circulating detached from its context. A widely repeated claim that the AI side solved at 3.2 times the rate, narrowing to 1.69 times among the top 5 per cent, comes from a different event, NeuroGrid CTF in November 2025, not from Nightfall. Attaching it to this year's benchmark would be wrong, and you will see it done.
Third, Hack The Box has an earlier and separate benchmark report from March 2026 with different figures entirely, covering 1,078 teams including 120 agentic AI teams. Two reports, similar names, different numbers. Cite the wrong one in a client deck and somebody will notice.
The event framing itself is straightforward enough: the Global Cyber Skills Benchmark is a competitive capture-the-flag run at scale, which is a proxy for offensive skill rather than a measure of real engagement work. Nobody's production estate is a CTF, and a challenge designed to be solvable is not a client network.
We are being this fussy for a reason. Most automated penetration testing claims currently circulating trace back through three or four hops to a vendor document nobody in the chain has read, and the numbers mutate at every hop. If you are going to put a figure in a proposal, be able to say which event it came from, who measured it and what it counted.
Independent research points the same way#
Independent research points the same way, from a different direction, which is the strongest thing you can say about any finding.
A paper published this year, Autonomous LLM Agents and CTFs: A Second Look, tested agents against 30 web-based capture-the-flag challenges spanning 14 vulnerability classes. Its framing is a direct challenge to the marketing: it notes that agents "are increasingly proposed to automate offensive security tasks, with recent studies reporting near human-level success rates", and then finds that both general-purpose and purpose-built agents hit the same barriers. A general coding agent matched the specialised architectures at 19 of 30 solved.
A second paper worth tracking, CTF-ABACUS, reports that trace-verified exploits account for only 62 to 87 per cent of recovered flags. Read that carefully, because it undercuts the flag statistic in the benchmark above: some proportion of any agent's flags are obtained without the exploit actually being demonstrated. A flag count is a weaker claim than it sounds.
Set the whole picture out and it is neither the revolution nor the damp squib. Agents are genuinely useful, measurably faster on the work they can do, and they stop short of finishing. The best AI-assisted teams reportedly worked three to four times faster, which is a real productivity claim and a completely different claim from replacement.
The pattern across both papers is convergence rather than acceleration. Purpose-built offensive agents and general-purpose coding agents land in roughly the same place, which suggests the constraint is not architecture but something more fundamental about the task. For anyone forecasting when automated penetration testing closes the gap, that is a more useful signal than any single score.
Anyone hunting a decisive result in this undergrowth will be disappointed, and that is itself the finding. Two independent measurements landing in the same place is stronger evidence than one dramatic number, and it points at a plateau rather than a cliff.
Capable enough to escape, not enough to finish#
Practitioners are not waiting for the marketing to settle, and the counter-evidence is arriving from the same week.
On 26 August a security researcher posted in r/cybersecurity about Trail of Bits testing whether a cyber-capable model could break out of a standard isolation environment. The post is blunt about the result.
Trail of Bits just tested if GPT-5.6-Cyber can escape a sandbox that we commonly use as an isolation mechanism. It did. Three times.
That is the honest tension in this whole category. The same technology that stops four challenges short of a human team in a competition is capable enough to break its own containment, and capable enough to be used offensively for real: Infosecurity reported a ransomware crew abusing a mainstream AI coding agent for reconnaissance and exploitation, with the underlying research documenting ten victims.
Capability and autonomy are different automated penetration testing products#
Hold those two findings together and the shape becomes clear. Capability is real and rising. Autonomous completion is not there. An agent that finds 32 of 36 things and cannot tell you which four it missed is a powerful assistant and a dangerous substitute, and the entire commercial question in automated penetration testing sits in that gap.
The industry has noticed. More than 150 technology and security firms (we counted 155 on OpenAI's own signatory list on 29 August, against the "nearly 130" reported at publication) signed a collective cyber defence pledge this week, published by OpenAI, asserting that AI-enabled attacks will escalate in the coming months. Whatever you make of the politics, it is quotable third-party authority and it cuts both ways: if the attack side is accelerating, so is the argument for keeping a human in the loop on the defence side.
The commercial reading is that automated penetration testing is currently a productivity technology wearing a replacement technology's marketing. Buyers who understand the difference will pay for the first. Buyers sold the second will be disappointed in about nine months, and they will remember who sold it to them.
A fox that can open the bin is not a fox that can run the kitchen. The distance between those two capabilities is where every serious offensive security services engagement currently lives, and it is where the money should be.
Five honest moves for anyone selling this#
Here is the part that concerns anyone selling security work rather than doing it.
For the last eighteen months the safe marketing move for penetration testing services has been to put AI on the homepage and let the buyer assume the rest. That move just acquired a public counter-number. A prospect can now ask what your agent scored, and "32 of 36" is a real answer that somebody else will quote at you if you do not quote it first.
There is an organic search dimension too, and it is immediate. Buyers now research automated penetration testing by asking an assistant, and assistants quote pages that state specific, sourced numbers. A page that says "AI-powered testing" is unquotable. A page that says "agents solved 32 of 36 at the 2026 benchmark, here is what we do about the four" is exactly the kind of passage a model lifts, with the source attached.
Put the benchmark result in your own materials before a prospect finds it. Owning an inconvenient statistic is the fastest available trust signal.
Your value is the four challenges the machine did not finish, and knowing which four they were. Price and describe that explicitly.
Three to four times faster is a real claim. Complete is a different claim. Never let a proposal blur them into one sentence.
Document who checks agent output, against what, and what happens when it is wrong. This is the deliverable buyers actually want.
Audit your site for unfalsifiable AI language. Anything a competitor could disprove with a public benchmark should go this week.
That last one is not decorative. Regulators are increasingly willing to act on what a vendor claims its automation does, and a claim you cannot evidence is a liability rather than a differentiator. Honest positioning is not a moral stance here, it is a risk control, and it happens to be the more persuasive pitch in a market full of identical assertions.
For penetration testing companies specifically, the differentiator is moving from finding to explaining. If agents narrow the gap on discovery, then the scarce good becomes the report a board can act on: what was found, what it means commercially, what to fix first and what to ignore. That is a writing and judgement problem, which is exactly where a cybersecurity marketing programme should be pointing.
The same applies to anyone selling pentest as a service on a subscription. Recurring revenue needs a recurring reason to exist, and "we run the scanner again" is not one when the scanner is commoditising. The reason has to be the human layer on top, described plainly enough that a chief financial officer can see what renews.
Nobody outfoxed anybody last week. A field of very good humans and a field of increasingly good machines ran the same course, and the humans finished it. That will not be true forever, and it is true now, which is the only ground a marketer is ever allowed to stand on. Build the positioning on the measured thing rather than the imagined one, and when the number changes, change with it. If you want help writing that honestly, come and talk to us, and read our content work first if you would rather see the method before the meeting.
One caution for the paid search side of this. Bidding on automated penetration testing terms while your landing page makes claims the public benchmark contradicts is an expensive way to lose credibility at speed, because the prospect arrives already primed by whatever the assistant told them on the way. Fix the page before you raise the budget. The research is public and your competitors can read it too.
Keep the scent on the evidence rather than the mood. Benchmarks will move, papers will follow, and the vendors who quoted the measured number this quarter will find it far easier to update their story next quarter than the ones who quoted a feeling. Automated penetration testing is going to keep getting better, and saying so plainly costs nothing.
Frequently asked questions#
Can AI replace penetration testing?
Not on current evidence. At the 2026 Global Cyber Skills Benchmark the best AI-assisted team solved 32 of 36 challenges while the team that solved all 36 was human. Agents are measurably faster on work they can do, and they stop short of completion, which is a different product from replacement.
How does automated penetration testing differ from a scanner?
A scanner checks a fixed list of known conditions. An agent chains steps, reasons about what it finds and attempts exploitation, which is closer to how a human tester works. The 2026 benchmark suggests agents do this competitively but stop short of finishing, so review remains part of the work.
How good are AI agents at capture the flag competitions?
Competitive but not dominant. Agent accounts made up 2.7% of registrations at the 2026 benchmark and produced 4.2% of flags and 4.6% of points, with 33 of the Top 100 teams including an agent. Independent research found a coding agent solved 19 of 30 web challenges.
Is automated penetration testing worth buying?
It is worth buying for speed and coverage on repeatable work, and worth questioning as a complete substitute. Ask any vendor what their agent scored on a public benchmark, who reviews its output, and what happens when it misses something. Those three answers separate the serious offers.
What should penetration testing companies say about AI now?
Name the benchmark number before a prospect finds it, sell the gap the machine leaves rather than the machine, keep speed claims separate from coverage claims, and document the human review step. Unfalsifiable AI language on a website is now a liability.
Does pentest as a service still make sense?
Yes, if the recurring value is the human layer rather than re-running a scanner. As discovery commoditises, the scarce good becomes the report a board can act on: what was found, what it means commercially, what to fix first and what can safely wait.
What should buyers ask offensive security services vendors?
Ask what their agents scored on a public benchmark, who reviews agent output and against what standard, how missed findings are caught, and whether speed claims and coverage claims are stated separately. Vendors who answer those cleanly are usually the ones doing the work properly.
Read more on this topic#
OpenAI's Own Agents Formed a Swarm
What happens when the agents are inside your estate rather than in a competition.
Read the pieceThe AI Malware Panic Was Loud. The Numbers Were Quiet.
The same discipline applied to the other big AI security claim of the month.
Read the pieceMarketing Built the Breach Surface at Three Airports
Where the real breaches actually happen, according to the regulator's own data.
Read the pieceCISA Tested Two SOCs. Only One Had an Incident Response Plan That Held.
Testing is only worth what the response plan behind it is worth.
Read the pieceSelling security services into a market full of identical claims?
folkfox writes positioning for cybersecurity firms that is specific enough to be checked. We start from the measured number, not the imagined one, because the measured one is the only thing that survives a procurement process.