Twelve websites.
Every one behind a WAF.

Every target in this programme was protected by a commercial web application firewall, in blocking mode. We confirmed that protection was live and enforcing — a control scan we ran first was detected and blocked. The real engagement was never blocked once.

12
Live websites
84
Findings reported
76
Distinct remediation tasks
39 min
Median per target
11h 27m
Total, reports included

We ran the test twice, on purpose

Any vendor can claim they got past your defences. The claim is worth nothing unless you can show the defences were working — so before the real engagement, we ran a control.

Control · commodity CVE sweep

WAF
ScannerYour app

SOURCE ADDRESS BLOCKED

The WAF recognised it and shut the address down. Exactly as designed — this is our proof the protection was live and enforcing.

The real engagement · original exploits

WAF
Rotating infraYour app

NEVER BLOCKED — 12 / 12 TARGETS

Nothing allowlisted, nothing recognised. There is no signature for an attack written during the engagement.

Control · commodity CVE scan

Blocked

An ordinary vulnerability sweep from a separate source address — the kind of automated scan that a great many firms sell as a penetration test. The WAF recognised it and blocked that address, exactly as it was designed to.

The real engagement

Never blocked. Not once.

Twelve targets, no allowlisting, original exploits written against each client's own code during the engagement. Not one of the twelve WAFs raised a block. There is no signature for an attack that did not exist before we started.

Your WAF is not the problem. It works.

It blocked the scan. That is the job, and it did it. The uncomfortable part is what that tells you about the testing market: a commodity scan is the thing your existing controls already stop. If that is what your last penetration test actually ran, you paid for a report on attacks that were never getting through in the first place.

The attacks that matter are the ones written specifically for your system, by someone asking what can be done to it — not looking up what has already been published about somebody else's.

What the swarm reached

On every one of the twelve sites, testing reached data or infrastructure that should not have been reachable — and proved it by retrieving it, with working code in the report.

Cloud credentials

Live AWS keys recovered from running systems — not found in a public repository, taken from the environment itself.

Personal data

Customer PII reachable through the application, proven by retrieval rather than inferred from a misconfiguration.

Operating system access

Shell and SSH access on hosts behind the application tier, established with working exploit code.

Read the shape, not the total

Anyone can generate a large number of findings. What tells you whether testing was real is how those findings are distributed.

Critical8 findings
High21 findings
Medium47 findings

What a scanner produces

A pyramid. Hundreds of informational and low-severity items, a handful of mediums, and almost nothing above that — because confirming a finding is genuinely critical means exploiting it, and automated tooling does not exploit anything.

What this programme produced

29 of 7638% — rated high or critical, and no informational padding at all. A top-heavy distribution is only possible when the findings were proven rather than matched. Several of those highs are critical once chained with others.

How long it took

Measured from starting a target to a finished report for that target. We publish the whole spread rather than the flattering average — one target took nearly three hours, and it got the time it needed.

12 min
Fastest target
39 min
Median target
52 min
Average target
2h 57m
Slowest target

The report is written in about three minutes

This is the number worth pausing on if you have bought a penetration test before. Testing finishes, and then you wait — five to ten business days is normal, by which time your engineers have moved on to something else and the urgency has drained out of the room. Here the report is part of the run, not a phase that starts after it. Twelve full targets, every report written, in eleven and a half hours.

Proven Attack Paths

Findings are scored one at a time. Attackers do not use them one at a time.

Every target in a programme feeds what it learns into the rest. A credential recovered from one site is tried against the others. An access path proven in one place informs how the next one is approached.

When all targets are finished, a final pass reads every report together and works out what an attacker could actually achieve by combining them — and how far the chains run. This is routinely where a set of individually medium findings, spread across separate systems, turns out to add up to one critical business outcome.

Risk does not add up. It compounds.

Why a traditional engagement cannot do this

Not for lack of skill — for lack of structure. Testers are assigned per target and work to a per-target time budget. Nobody is paid to sit down at the end and correlate twelve separate reports against each other, so nobody does.

And this is not modelled from configuration data the way attack-path graphing tools do it. These chains were walked, using credentials actually captured and exploits actually proven.

Why the results look like this

“We look at this like a programmer doing a code audit and a red team engineer: what can I do to this? Not what did other people find. I don't care what they found.”

That distinction is the whole difference between the two columns above. Testing that asks what has already been published about other people's software is a lookup, and your existing controls already handle it. Testing that asks what can be done to your code produces things nothing has a signature for yet.

The swarm was built against one acceptance criterion, applied by someone who has been doing this for thirty years: can it do this better than I can? Where the answer was no, it was not shipped — it was improved until the answer changed.

The part nobody talks about: the weeks after

A penetration test report is an uncomfortable object. It is a written, evidenced list of ways into your business, and closing those findings takes your engineering team weeks and a release cycle. Between delivery and remediation you are not less exposed — you are exposed in writing.

So protection during that window is part of the engagement, not a separate product sold to you afterwards. Monitoring is configured directly from the confirmed findings — it knows the exact request that worked against your system, because it is the one we sent — and it can block that attack across every protected host while your team fixes the root cause. Coverage runs until the finding is retested and closed. Then it is your call whether it stays.

How an engagement works

Questions about these results

Find out what is actually reachable

Price your scope in about a minute, or read the full unedited report first — no email, no form, no follow-up sequence.