For the person buying the test

Three things get sold as a penetration test.

They cost about the same and arrive looking about the same. APEX exists so you can tell which one you are buying before you pay for it.

We prove what you can do. So clients can trust what you say.

01

A scan with a cover letter

A CVE scan. A tool is pointed at the infrastructure by somebody who was told to point it there, the output is tidied into a template, and it is sent over. It reports what is already published about software you are running. Nobody attempted to exploit anything, so nothing in it was ever proven — and by construction it can only ever find what somebody else already found and wrote up. It will not once find something unpublished, because it is not looking. It is sold as a penetration test and priced like one.

Ask which findings they exploited. Ask to watch one.

02

A real test — done the old way

Genuine work by somebody who knows what they are doing. A methodology, actual exploitation, findings that hold up. Nothing wrong with any of it except the arithmetic: one person, a fixed window, and far more endpoints, parameters and code paths than a human can work through by hand. So they pick, and picking means deciding up front what to leave untouched — and they will not know what was in the part they skipped. That is not a criticism of the tester. It is a limit that no longer has to be accepted, because AI can cover that volume while the engineer supplies the judgment.

Ask what they did not get to, and why. Then ask what they would have found with ten times the coverage.

03

An engineer who could do the second one, amplified

Everything above, plus coverage no individual could reach unaided — every request, every parameter, every reachable path — because the model does the volume and the engineer supplies the judgment about what matters. Your data stays inside approved systems while it happens, and every finding still arrives with code somebody else ran.

This is what the credential certifies.

The three arrive looking similar and cost similar money. The credential exists so you can tell which one you are getting before you pay for it, instead of finding out when somebody else gets in.

What the credential guarantees

“You have that cert?” Then these are true.

Holders don’t need AI to do the job. They know how to use it to do more than they could alone.

The engineer is doing the work, not the model.

One exam phase runs with generative AI disabled entirely, and it is a pass/fail gate no score elsewhere can compensate for. The oral defense also runs with the models off — the board picks lines out of their code and their report and asks them to explain it, live, unprepared.

Unplug the AI and they can still run your engagement. It takes longer, and more of your infrastructure goes unread — which is the honest reason the good ones use it.

They get through your whole infrastructure, not a sample of it.

An engagement produces more data than a person can read — every request, every parameter, every reachable path, every line of code in scope. Working by hand, an engineer picks a subset and the rest goes untested, which is a limit of arithmetic and not of skill. The exam measures whether the candidate can drive AI across the entire surface and still apply judgment to what comes back, because coverage without judgment is just a longer report.

Without AI, somebody decides what to skip and you never hear which parts. With it used properly, far less gets skipped — and the report tells you what did.

They understand your infrastructure, not just a methodology.

Before they are allowed to attack anything, they have to build a full enterprise infrastructure themselves — two clouds, on-premise virtualization, a routed core, software-defined networking, the CDN and edge in front of it, remote access, perimeter, a mixed Windows, Linux and macOS fleet, managed mobile devices and a working voice platform — and defend every design decision they made.

You cannot break what you cannot build. They built it.

Your data does not end up in somebody’s chatbot.

The issue was never whether a model was used — it is where the data went. A private or self-hosted deployment under your contract raises none of this; a consumer chatbot raises all of it, and the same engagement data is worth the same to an attacker either way. Candidates build the controls that enforce that distinction themselves — classification, outbound inspection, per-engagement isolation, canaries, fail-closed routing — and then those controls are attacked. A confirmed leak of protected data ends the exam attempt on the spot, whatever the rest of the score says.

There is finally a credential that fails people for the thing a lot of testers are quietly doing right now.

They find what a scan never will.

The exam requires an original, non-public issue derived from code, architecture, state or runtime behavior. A scanner result, a version match or a copied exploit is explicitly not accepted as a finding. Where a known vulnerability is present it gets documented — but the exploitation is original, because a published proof only establishes that a version matched, not that it is genuinely reachable past your controls in your configuration.

We do not take the easy way into your network. A CVE scan can only ever find what somebody else already published.

Prove it — every finding has a working exploit behind it.

Not "this could be exploited." An engineer wrote a proof of concept, a reviewer took it into a clean environment and ran it, and it visibly crossed the boundary being claimed. Scanners and unsupervised AI hand you "could be done" — a list of theories with severity ratings attached. That is not a test.

Put your skills behind your words. And if an APEX holder could not exploit it, very few people or systems were going to.

The report tells you what it did not test.

Coverage is accounted for explicitly: every surface is mapped either to evidence or to a stated limitation. Missing a designated critical issue is a fail even when every submitted finding is accurate. Silence about coverage is scored as a gap, not as completeness.

A clean report that quietly skipped half your infrastructure is the thing that gets you breached.

And when they could not

A negative result from them means something too.

It also means something when they could NOT exploit it. An engineer who has proven they can drive AI across an entire attack surface and still could not cross the boundary has told you something a scanner never can: it was genuinely covered, and it held.

The missing link

The problem is not that engineers use AI. It is that most use it badly.

The public argument is still stuck on whether generative AI belongs in security testing. That argument is a distraction, because the people causing the damage are not waiting for it to be settled. They are already using it, with no structure around it, and the results are going out to clients.

Point a model at an engagement with nothing enforcing rigor and it will produce volume — plausible findings, confident severity ratings, remediation advice, and a substantial amount that is simply not true. An engineer who cannot tell the difference passes it straight through. The client gets a longer report and a worse test.

That is where the false positives come from, and it is why so many buyers have concluded that AI-assisted testing is noise. They are drawing a fair conclusion from real evidence. They have just never been shown the other version.

Because used correctly, it is the largest coverage gain this work has ever had, with an experienced engineer deciding what any of it means. Used incorrectly, it is a false-positive generator with excellent grammar. Nothing on the outside of a report tells you which one you paid for.

That is the gap this credential closes. Not whether they used AI — whether they are any good at it.

The honest part

What the credential does not promise.

Two things worth being straight about. A certification that oversells itself is worth less than one that does not, and you should be able to tell where this one stops.

It is still not a guarantee.

A qualified engineer can miss something — anyone telling you otherwise is selling you something. The credential establishes that the person is capable, that your data was governed, and that whatever they reported was proven with code somebody else ran. Not that your infrastructure is now safe.

It says nothing about the firm.

It certifies a person, not their employer, and it is deliberately open to engineers at firms that compete with us. Scope, contract, insurance and how the firm handles your data at an organizational level are separate questions. Still ask them.

Use this on anyone

Six questions worth asking any tester.

Ours or somebody else’s. If a vendor answers these well, that tells you more than any logo on a proposal — including this one.
  1. Which AI did you use on my engagement, and what did it see?

    A confident answer names the provider, the deployment, and the data classes it was allowed to touch. "We use AI to help" is not an answer.

  2. Can you show me the log of what was sent to it?

    If there is no log, nobody knows what left your environment — including them.

  3. Which findings did you personally reproduce by hand?

    The honest answer is a number. If it is "all of them" on a large report, ask how long that took.

  4. What did you not test, and why?

    Every real engagement has gaps. A tester who claims none either did not look or is not telling you.

  5. Can you show me the proof-of-concept code, and will you run it in front of me?

    This is the one that separates a report from a document. If the "proof of concept" is a screenshot and a paragraph, you were told a vulnerability exists — you were not shown one.

  6. Could you have found this without AI?

    Not a gotcha. You are asking whether the tool amplified an expert or replaced one.

We are handing you these knowing some get asked of us. That is the point — a firm that would rather you did not ask is telling you something.

Checking it

Name and number, together.

Verification needs the holder’s name and their credential number. A number on its own returns nothing — the registry confirms a person you are already looking at rather than acting as a browsable directory of certified engineers. Status is live, so a suspended or revoked credential shows as such immediately, whatever the printed certificate says.

One thing worth doing today

Find out where the source code and screenshots from your last penetration test ended up. Most buyers have never asked, and most contracts do not cover it — because when those contracts were written, the question did not exist.

More detail on the standard itself: the evidence standard and the data-protection standard.