SA-APEX · performance exam

Six phases. Nothing you can answer from memory.

There is no multiple choice, no attendance credit and no pass for a well-formatted report. Every phase produces an artifact that a reviewer can pick up, run, and challenge you on.

Phase by phase

What each phase measures, and what the models are allowed to do.

Phase A: Manual foundation

AI disabledWaived on a verified OSEE bridge

Do the work with the models switched off.

  • No generative AI of any kind. The candidate works from source, protocol behaviour and their own tooling.
  • This phase exists to establish that everything measured later is amplification of real skill rather than a substitute for it.
  • A candidate who cannot complete or explain this work does not proceed, regardless of how strong the remaining phases are.

Deliverable Manual exploitation, code modification, attack-path notes, live spot checks.

Phase B: Build the estate

Allowed through the controlled gateway

Design, build and configure a full enterprise estate from bare metal up.

  • The candidate builds the kind of environment they claim to be able to attack: cloud and on-premise virtualisation, a routed core, remote access, perimeter and endpoint control, a mixed operating-system fleet, mobile devices and a working voice platform.
  • The build is graded on whether it stands up and whether the candidate can explain every design decision — not on how hard it is to break.
  • A build that does not pass the conformance run does not enter the target pool, and the candidate does not proceed to the attack phase on it.

Deliverable A running estate, its configuration as code, a design record, and a build conformance run that passes.

Phase C: Build the tooling

Approved sandbox only

A bare Kali server and an AI API token. Nothing else.

  • The candidate is given a clean Kali server and a single AI API token. No agent framework, no prepared prompts, no Security Arsenal tooling.
  • They build the thing that does the testing: scoped agents, typed tools, an enforced provider policy, evidence capture and a kill switch.
  • They also build the controls that stop engagement data reaching the model — classification, redaction, tenant isolation, canaries and outbound inspection — and then prove those controls hold under an adversarial test.
  • A thin wrapper around a chat API does not pass this phase.

Deliverable A candidate-built AI security application, its threat model, its policy enforcement and its test suite.

Phase D: Attack a different estate

Allowed through the controlled gateway

You built one. Now break one you have never seen.

  • The estate assigned for the attack phase is never the one the candidate built. It comes from the graded target pool, anonymised and randomly assigned.
  • Findings must be derived from architecture, code, state, data flow, trust boundaries or runtime behaviour — not from a disclosed CVE, a version banner or a scanner verdict.
  • Scored defects are seeded after a build passes conformance, so every candidate is measured against the same objective set regardless of whose estate they drew.
  • Unseeded attack surface is deliberately left in place. Genuine discoveries outside the seeded set earn expert review.

Deliverable Original findings, proofs of concept, an evidence graph and a coverage ledger.

Phase E: Reporting

Allowed, fully logged

Evidence-locked reports, plus a full account of what you did with AI.

  • Report generation reads only findings that have passed the evidence gates. The generator cannot invent facts outside locked fields.
  • Every severity, affected asset, impact statement and reproduction claim carries an evidence link back to an immutable artifact.
  • The coverage ledger has to name what was not tested and what uncertainty remains. Silence about coverage is scored as a gap, not as completeness.

Deliverable Executive and technical reports, an evidence package, a coverage ledger, an AI-use manifest and a validation manifest.

Phase F: Oral defense

AI disabled

Explain it live, with the models off.

  • The board samples code, commands, findings and report language unpredictably. The candidate owns every line they submitted.
  • AI use during the exam is expected and logged. Hidden AI use is not the problem being tested for — an inability to explain the result is.
  • Contradicting your own submitted evidence ends the attempt.

Deliverable Live explanation, replay, challenge questions, remediation defense.

On exam length

The exam runs as scheduled blocks across a window rather than a single sitting — the build phase alone does not compress into an afternoon. Exact durations are being set from pilot runs. We will publish the number when we have measured it, and not before.

Scoring

One gate, six weighted domains, and failures that cannot be averaged away.

Competency domains and weights
DomainWeightWhat mastery looks likeNon-compensable failure
AManual offensive foundationPass/fail gateCan scope, test, exploit, pivot, reason and report with generative AI switched off.Any inability to complete or explain core work without AI.
BBuild and operate the estate15%Designs, builds and configures the infrastructure being attacked — cloud, virtualisation, routing, VPN, firewall, endpoints, servers, workstations, mobile and voice — and can defend the design choices.A build that does not stand up, or a design the candidate cannot explain.
CSecure AI architecture and data governance20%Selects safe deployment patterns, enforces data boundaries and prevents leakage.Protected-data leak, bypassed gateway, or an unapproved provider route.
DAI application and agent engineering15%Builds reliable tools, sandboxes execution, controls permissions and logs provenance.Unbounded tool access, missing tenant isolation, or a build that cannot be reproduced.
ENovel discovery and exploitation20%Finds non-public logic and security defects and creates safe proof of exploitability.No original issue proven; exploit not understood; no security boundary crossed.
FEvidence and false-positive control15%Uses independent evidence, replay, negative tests and calibrated confidence.A material false positive in the final report, or fabricated evidence.
GReporting and coverage assurance15%Produces traceable reports, supports every statement and accounts for missed critical risk.A designated critical issue missed, or an unsupported executive claim.

The manual gate is not a weighting.

Domain A is pass/fail. A candidate cannot average past a failure in manual capability with a strong AI build, and no combination of the other six domains compensates for it. Everything else in SA-APEX assumes that gate held.

Why the thresholds are not published.

A pass mark that has not been through standard setting is a guess with a number on it. Publishing it would invite candidates to optimise against a figure we have not yet validated. The domain minimums, the gate and the critical-miss rule are real and are described here. The numbers arrive with the pilot data.

AI during the exam

Expected, permitted, and logged in complete detail.

Candidates work through a controlled examination gateway. It records prompts, model and version, retrieved context, tool calls, outputs, policy decisions and cost. Model availability is standardised per exam form so two candidates are not scored against different capabilities.

You may build local components. External network and model calls are blocked unless the phase explicitly provides them.

The score rewards the engineering and the verified result. It does not reward token consumption or agent count.

Integrity and authorship

  • Included:You own and must be able to explain every submitted line of code, command, finding and recommendation.
  • Included:Hidden AI use is not the problem being tested for. Misrepresentation, outside assistance, answer sharing and unapproved endpoints are.
  • Included:The oral defense samples your work unpredictably, which is what distinguishes genuine ownership from generated output.
  • Included:The platform retains enough provenance to support an appeal without publishing exam secrets.

Immediate fail conditions

These end the attempt, whatever the rest of the score says.

  • Excluded:Using AI to explain code the candidate cannot independently explain during oral defense.
  • Excluded:Uploading protected exam or client-like data to an unapproved endpoint or a public model.
  • Excluded:Acting outside scope, disabling safety controls, attacking shared infrastructure, or reaching another candidate’s work.
  • Excluded:Claiming scanner output, a version match or a model assertion as an exploitable finding without proof.
  • Excluded:Fabricating evidence, hiding a failed validation attempt, altering timestamps, or misrepresenting model or tool actions.
  • Excluded:Failing to identify or report a planted critical issue that the approved coverage map makes reasonably discoverable.
  • Excluded:Producing a materially unsafe exploit or remediation recommendation that would predictably harm the controlled target.

Five of these seven are things that would harm a real client. That is not a coincidence — the exam is calibrated against consequences, not against difficulty for its own sake.

After the exam

Grading is separated from teaching, and from you.

Assessment roles
RoleWhat they can do
ProctorSession controls and integrity events. No unilateral scoring changes.
GraderAssigned anonymised submissions, the rubric, and a replay environment. They see a candidate handle, not a name or an employer.
Senior reviewerCritical findings, legitimate alternate solutions, score exceptions and the oral defense.
Appeals panelThe appeal record and the evidence needed to decide it, independent of the original decision wherever practical.
Program adminCatalog, policies, scheduling and credential lifecycle. Restricted access to exam content, and no silent score edits — every change carries a reason, an authorization and an audit row.

Anonymised grading matters more here than in most programs, because SA-APEX is open to people who work for firms that compete with us. A reviewer who can see the employer is a reviewer whose decision can be questioned. Appeals and retests.