Security Arsenal performance certification

Master the attack before commanding the swarm.

SA-APEX — the Security Arsenal Advanced Penetration Engineering Expert. A certification for engineers who already do this work, and who now have to do it with AI in the room without losing confidentiality, accuracy or accountability.
  • You build the estate
  • You attack somebody else’s
  • Two phases with AI switched off
  • Defended live in front of a board

The name

APEX stands for something. Here it is.

Advanced

You arrive able to do the work. This measures the ceiling, not the floor.

Penetration

Real offensive testing against real infrastructure. Not theory, not a quiz.

Engineering

You build the estate and you build the tooling. An operator who cannot build is guessing.

Expert

Proven in front of a board that can interrupt you, with the models switched off.

Advanced Penetration Engineering Expert. The credential is named for the person, not for a model or a vendor, because the models will change three times before the credential does — and the thing being certified is the engineer either way.

The origin

Why this exists.

Thirty years into doing this work, something broke. Generative AI got good enough that an underqualified person could produce a penetration test report that reads better than a real one — clean severity ratings, confident prose, tidy remediation. Everything except the part where somebody actually got in.

Presentation quality used to be a reasonable proxy for testing quality. It is not any more. A client can now receive a beautiful document and still be wide open on the path that actually matters, and they have no way to tell from the outside. They pay either way. They carry the breach either way.

The industry response has mostly been to argue about whether AI belongs in security testing at all. That is the wrong argument. It is here, it makes a good engineer substantially more dangerous, and it makes a bad one completely unaccountable. The question worth answering is: how do you tell which one you hired?

You cannot tell from the report. So we certify the work underneath it — whether the engineer can do the engagement without AI at all, whether they can build the infrastructure they claim to be able to break, whether client data stayed inside approved systems, whether every finding survives somebody else reproducing it, and whether they can defend the whole thing live with the models turned off.

That is the entire product. Everything on this page is a consequence of it.

The methodology it comes from

“We look at this like a programmer doing a code audit and a red team engineer — what can I do to this? Not what did other people find. I don’t care what they found.”

Original exploits against the client’s actual code. Everything proven with a working proof of concept. Nothing guessed. The swarm was built by the person who still does engagements — the best part of the job — against one criterion: can it do this better than me? If not, make it better. That is the bar this exam is set to.

Before you go any further

This is not a penetration-testing course. If you cannot run the engagement, write the tooling, prove the exploit and defend the report with every model switched off, you are not ready for this yet.

We would rather cost you five minutes on a free page than a failed exam. There is a self-check that runs entirely in your browser and is perfectly willing to tell you no.

What it proves

Four things, each proven by an artifact rather than an answer sheet.

01

Manual mastery

The engineer can execute a complete engagement without generative AI.

Proven by AI-disabled practical phase, oral defense, code comprehension checks.

02

Safe AI engineering

Client data stays inside approved boundaries and every tool has deliberate controls.

Proven by Architecture submission, data-flow review, leakage tests, policy enforcement.

03

Novel discovery

The engineer can find and prove weaknesses that are absent from scanner and CVE output.

Proven by Seeded and unseeded defects, exploit construction, impact evidence.

04

Defensible truth

Findings and reports are evidence-based, complete, reproducible and independently replayed.

Proven by Evidence graph, independent replay, coverage accounting, report defense.

The structure

You can’t break what you can’t build.

Most certifications hand you a target somebody else configured. That measures one skill and hides the one that matters — whether you actually understand the technology well enough to have built it. SA-APEX inverts it.

Step one

You build the estate.

Cloud and on-premise virtualisation, a routed core, remote access, perimeter, a mixed operating-system fleet, managed mobile devices and a working voice platform. You design it, you configure it, you defend the design.

Step two

You build the tooling.

A clean Kali server and one AI API token. No agent framework, no prepared prompts, no Security Arsenal software. You build the thing that does the testing — and the controls that stop client data reaching the model while it works.

Step three

You attack a different one.

Never your own build. You draw an anonymised estate from the graded pool and go at infrastructure a real engineer configured — with all the inconsistency that implies.

What you have to be able to stand up

Cloud and virtualisation

  • AWS
  • VMware
  • Proxmox

Network

  • BGP
  • OSPF
  • VPN
  • Firewall

Endpoints and servers

  • Windows
  • Linux
  • macOS
  • Endpoint security

Mobile and voice

  • Apple mobile
  • Android
  • VoIP

Every component is graded on whether it works and whether you can explain why you built it that way — not on how hard it is to break. Details on the build-and-break page.

What it is not

The exclusions are the product.

  • Excluded:Not beginner penetration-testing instruction, and not a "become a hacker" program.
  • Excluded:Not a certification for running a scanner, pasting the output into a model, or matching software versions to CVEs.
  • Excluded:No pass based on multiple-choice questions, attendance, course completion, or an AI-generated report.
  • Excluded:No uncontrolled use of consumer or public AI services with client code, credentials, screenshots, logs or findings.
  • Excluded:No finding accepted without a reproducible attack path, evidence, impact analysis and source or runtime reasoning.

This is not a credential about attacking AI systems.

There are good credentials that certify testing AI systems themselves — model behaviour, prompt handling, agent abuse, the AI attack surface. SA-APEX measures the other direction: an engineer using AI to test everything else, on real client work, without losing confidentiality, technical accuracy or personal accountability. The two do not overlap much, and holding one says nothing about the other.

The performance exam

Six phases. The models are off for two of them.

AI use inside the exam is expected and logged in full. What is being measured is whether you can work without it, engineer with it, and account for every claim either way.
  1. Phase A: Manual foundation

    AI disabled

    Do the work with the models switched off.

    Deliverable Manual exploitation, code modification, attack-path notes, live spot checks.

  2. Phase B: Build the estate

    Allowed through the controlled gateway

    Design, build and configure a full enterprise estate from bare metal up.

    Deliverable A running estate, its configuration as code, a design record, and a build conformance run that passes.

  3. Phase C: Build the tooling

    Approved sandbox only

    A bare Kali server and an AI API token. Nothing else.

    Deliverable A candidate-built AI security application, its threat model, its policy enforcement and its test suite.

  4. Phase D: Attack a different estate

    Allowed through the controlled gateway

    You built one. Now break one you have never seen.

    Deliverable Original findings, proofs of concept, an evidence graph and a coverage ledger.

  5. Phase E: Reporting

    Allowed, fully logged

    Evidence-locked reports, plus a full account of what you did with AI.

    Deliverable Executive and technical reports, an evidence package, a coverage ledger, an AI-use manifest and a validation manifest.

  6. Phase F: Oral defense

    AI disabled

    Explain it live, with the models off.

    Deliverable Live explanation, replay, challenge questions, remediation defense.

Exam length is being set from pilot data rather than from a marketing page. When we have run enough of them to know, we will publish the number. Full exam structure.

Recognised prior work

Hold an OSEE? One phase comes off. The rest does not.

OSEE (EXP-401Advanced Windows Exploitation) already proves manual exploitation at the depth our manual gate is set to. We do not need you to prove it twice.

Waived

Phase A — Manual foundation

You have already demonstrated it, under exam conditions, to a standard we respect.

Still required

  • Building the estate
  • Voice
  • Software development
  • Every AI phase

Security Arsenal recognises OSEE as evidence for our manual foundation gate. This is our own admissions decision. It is not a partnership, endorsement, affiliation or articulation agreement with OffSec, and it confers nothing on their behalf. How the bridge route works.

Client-data protection

A penetration test contains exactly what an attacker wants.

Source code, credentials, architecture, weaknesses, exploit paths, screenshots, logs, business context. The candidate has to prove all of it stays inside approved, isolated systems — and then prove it under an adversarial test.
Data classification and permitted AI destinations
ClassExamplesDefault AI destination
PublicPublished domains, public source, public documentation.Approved enterprise or private model. A public model only where company policy explicitly allows it.
InternalGeneric methods, sanitized templates, non-client operational notes.Approved enterprise tenant with contractual controls.
Client confidentialArchitecture, code, findings, screenshots, logs, asset names.Security Arsenal-controlled or client-approved private deployment. No consumer or public endpoint.
RestrictedCredentials, tokens, PII and PHI, private keys, production data, regulated records.Prefer no model ingestion at all. If essential: isolated approved deployment, minimized, encrypted, audited, explicitly authorized.
ProhibitedBlockedData barred by contract or law, another client’s data, unapproved secrets, exam answer material.Never submitted. Blocked and alerted.

Candidates build and pass six leakage tests, including cross-tenant retrieval, canary strings and prompt-injection exfiltration. The full control standard.

Low false positives by design

A fluent explanation is not a finding.

Generative AI can help an underqualified person produce a polished report that reads as authoritative and leaves the most important attack path untouched. Presentation quality stopped being evidence of testing quality. So we certify the work behind the report.
Finding state machine
StateDefinitionIn the report?
HypothesisA potential issue proposed by a human, a model, an analyzer or a test.Not in the report
CandidateCode or runtime evidence suggests a plausible attack path.Not in the report
ValidatedIndependent reproduction crosses the claimed security boundary.In the report
Conditionally validatedExploitability is proven under clearly stated conditions; constraints remain.In the report, conditions stated up front
RejectedNegative testing disproved the claim, or the evidence was insufficient.Not in the report; retained internally for scoring
Duplicate or variantSame root cause or attack primitive as another finding.Merged; affected paths preserved

Key doctrine

AI confidence is not evidence. Multiple agents repeating the same unsupported claim are not independent confirmation. A finding exists only when the candidate can show the affected asset and code path, the preconditions, the controlled action, the observed result, the security boundary crossed, the impact, and a repeatable method that reproduces it.

Beyond CVEs

Original, or it does not count.

What qualifies as an original discovery
QualifiesDoes not
A business-logic or authorization flaw derived from code and runtime behaviour.A scanner saying IDOR or BOLA may exist.
An unsafe interaction between components that are individually reasonable.A dependency with a known CVE.
A state-machine, race, cache, tenancy, trust-boundary, parser or workflow weakness.A generic missing-header observation.
A candidate-built proof crossing a security boundary with controlled impact.Model-generated exploit text that was never understood or executed.
Variant analysis showing related reachable paths and a shared root cause.One copied request with no explanation.

Version and CVE scanning is useful inventory work. It is not the skill being measured here, and no amount of it substitutes for reading unfamiliar code until you can predict where it breaks.

The journey

Course completion is not certification.

  1. 1Apply

    Submit experience, languages, build and engagement history, and identity.

    Outcome Eligibility review. Nothing is sold before prerequisites are clear.

  2. 2Baseline challenge

    Complete a short manual coding and security reasoning screen.

    Outcome Admit, require specific remediation, or decline.

  3. 3Security AI Engineering

    Work through privacy architecture, model and tool engineering, evidence discipline, orchestration and reporting.

    Outcome Course completion record only — not a certification.

  4. 4Practice range

    Build and harden AI testing applications, and complete replayable scenarios.

    Outcome Readiness signals and an evidence portfolio.

  5. 5Exam readiness

    Meet the objective readiness gates and self-attest.

    Outcome Performance exam scheduled.

  6. 6SA-APEX exam

    Build, tool, attack, report and defend.

    Outcome Pass, fail, or a limited retest decision.

  7. 7Credential

    Accept the code of conduct and set publication preferences.

    Outcome Verifiable credential issued.

  8. 8Internal authorization

    Security Arsenal employees only: complete company controls and supervised operation.

    Outcome Revocable SwarmPT Lead Authorization.

Who it is for

We built this for the whole profession. Including our competitors.

When an AI-assisted test is inaccurate or incomplete, the client carries the risk. Not the firm that wrote the report. Not the engineer who ran it.

That is a bad arrangement, and it does not get better if the standard of care only exists inside one company. So SA-APEX is open to any qualified practitioner, whoever they work for — including firms that compete with us directly, for the same clients, in the same week.

Grading is anonymised. A reviewer sees a candidate handle, never a name or an employer, so a decision about a competitor’s engineer cannot be questioned and cannot be leaned on.

We would rather compete against engineers who are held to this standard than win against ones who are not.

Working practitioners

Testers, red teamers, application security engineers, exploit developers and senior security engineers who already run authorized engagements.

Internal security teams

People who test their own estate and need a defensible standard for what may go into a model and how a finding is established.

Testing firms

Including ours. Our engineers sit the same exam, graded the same way, against the same rubric, with the same handle instead of a name.

What we have not decided yet

We publish prices for everything else we sell. We will publish these too — when they are real.

Exam length, pass standard, price, credential term and proctoring model are being set from pilot evidence and a standard-setting panel. Presenting a design target as a settled number would be the first dishonest thing this program did, so we are not doing it.

Straight answers

The questions people actually ask.

What is SA-APEX?
SA-APEX (Security Arsenal Advanced Penetration Engineering Expert) is a performance-based certification for experienced penetration engineers who use AI on authorized client work. It measures manual capability with AI switched off, infrastructure engineering, secure AI architecture, original code-derived vulnerability discovery, evidence-based reporting and a live oral defense.
Is SA-APEX a certification about hacking AI systems?
No. There are separate credentials that certify testing AI systems themselves — model behaviour, prompt handling, agent abuse. SA-APEX measures the other direction: an engineer using AI to test everything else, without losing confidentiality, technical accuracy or personal accountability. Holding one says nothing about the other.
Is SA-APEX suitable for beginners?
No. It assumes the candidate can already conduct a complete engagement, develop tooling, prove an exploit and defend a report without any generative AI. A candidate who cannot do that does not pass the manual foundation gate, which no score elsewhere compensates for.
Why does the SA-APEX exam make you build a network before attacking one?
Because you cannot break what you cannot build. The candidate designs and configures a full enterprise estate — cloud and on-premise virtualisation, BGP and OSPF routing, VPN, firewall, Windows, Linux and macOS fleets, managed Apple and Android devices, and a working VoIP platform — and is then assigned a different candidate’s estate to attack. Building it proves the understanding that tells an engineer where to look.
Does an OSEE waive part of the SA-APEX exam?
An active OSEE (OffSec EXP-401, Advanced Windows Exploitation) satisfies the manual foundation gate and waives that phase only, once the certification number is verified with the issuing body. It does not waive the build phase, the VoIP requirement, the software development requirement or any AI phase. This is Security Arsenal’s own admissions decision, not a partnership or endorsement.
Can engineers at competing security firms earn SA-APEX?
Yes, deliberately. When an AI-assisted test is inaccurate or incomplete the client carries the risk, and that does not improve if the standard of care exists inside only one company. SwarmPT remains private to Security Arsenal employees; the credential conveys no product access and no authorization to operate on any client system.
Does completing the Security AI Engineering course make you SA-APEX certified?
No. Course completion produces a separate record with its own identifier and expiry, marked as not being the certification. Certification comes only from passing the independent performance exam, and many people who complete the course will not pass it.
How do you verify an SA-APEX credential?
At securityarsenal.com/verify, using both the holder’s name and their credential number. Both are required: a credential number on its own returns nothing, so the registry confirms what a verifier already knows rather than acting as a browsable directory of certified engineers. Verification shows the current status — active, expired, suspended or revoked — never exam scores or employer.
What does the SA-APEX exam cost and how long does it take?
Neither is published yet. Exam length, pass standard, price and credential term are being set from pilot evidence and a standard-setting panel rather than presented as settled numbers. Security Arsenal publishes the open decision log at securityarsenal.com/certifications/sa-apex/decisions.
What ends an SA-APEX attempt immediately?
Using AI to explain code the candidate cannot independently explain; uploading protected data to an unapproved endpoint or a public model; acting outside scope or attacking shared infrastructure; claiming scanner or model output as an exploitable finding without proof; fabricating evidence; missing a designated critical issue that the coverage map makes reasonably discoverable; or producing a materially unsafe exploit or remediation.