Security Arsenal performance certification
Master the attack
before commanding the swarm.

- You build the infrastructure
- You attack somebody else’s
- Two phases with AI switched off
- Defended live in front of a board
Models off · one engineer, by hand
0 of 48 paths tested
Covering the surface is what the model is for. Deciding what any of it means is not. The exam measures both, separately.
We prove what you can do. So clients can trust what you say.
The name
APEX stands for something. Here it is.
Advanced
You arrive able to do the work. This measures the ceiling, not the floor.
Penetration
Real offensive testing against real infrastructure. Not theory, not a quiz.
Engineering
You build the infrastructure and you build the tooling. An operator who cannot build is guessing.
Expert
Proven in front of a board that can interrupt you, with the models switched off.
Advanced Penetration Engineering Expert. The credential is named for the person, not for a model or a vendor, because the models will change three times before the credential does — and the thing being certified is the engineer either way.
The origin
Why this exists.
Thirty years into doing this work, something broke. Generative AI got good enough that an underqualified person could produce a penetration test report that reads better than a real one — clean severity ratings, confident prose, tidy remediation. Everything except the part where somebody actually got in.
Presentation quality used to be a reasonable proxy for testing quality. It is not any more. A client can now receive a beautiful document and still be wide open on the path that actually matters, and they have no way to tell from the outside. They pay either way. They carry the breach either way.
The industry response has mostly been to argue about whether AI belongs in security testing at all. That is the wrong argument. It is here, it makes a good engineer substantially more dangerous, and it makes a bad one completely unaccountable. The question worth answering is: how do you tell which one you hired?
You cannot tell from the report. So we certify the work underneath it — whether the engineer can do the engagement without AI at all, whether they can build the infrastructure they claim to be able to break, whether client data stayed inside approved systems, whether every finding survives somebody else reproducing it, and whether they can defend the whole thing live with the models turned off.
That is the entire product. Everything on this page is a consequence of it.
The methodology it comes from
“We look at this like a programmer doing a code audit and a red team engineer — what can I do to this? Not what did other people find. I don’t care what they found.”
Original exploits against the client’s actual code. Everything proven with a working proof of concept. Nothing guessed. The swarm was built by the person who still does engagements — the best part of the job — against one criterion: can it do this better than me? If not, make it better. That is the bar this exam is set to.
The obvious objection
There are already a hundred security certifications. Why another one?
Because every one of them certifies the job as it existed before a model was sitting next to the engineer. That is not a criticism — they measure real skill, and some of them are hard. It is that none of them ask the two questions that now decide whether a test was any good.
The first: can this person still do the work when the model is switched off? AI raised the floor on how a report looks. It did not raise the floor on how good the test underneath it was. An underqualified tester can now produce a document with clean severity ratings, confident prose and tidy remediation, containing nothing that anybody actually proved. You cannot tell from the outside. You pay either way.
The second, and the one nobody is testing for: where did your data go? A penetration test contains source code, credentials, architecture, screenshots, logs, and a list of exactly how to get in. Right now, across this industry, a lot of that is being pasted into consumer chatbots because it is faster. Every existing certification can be held, in good standing, by someone doing that today. Not one of them would notice.
That is the gap. Not another credential — the first one that measures what changed.
What we are actually worried about.
This is already happening without AI. A large share of what is sold as penetration testing today is a scan — run by somebody who was told to run it, with the output tidied up and sent over. Nobody in that chain ever attempted to exploit anything. The client receives a document, believes they were tested, and was not.
AI is about to make that far easier to sell and far harder to spot. Point the right setup at a network and something comes out the other end — findings, severities, remediation, a clean report. It will read better than the scan did. It will look like a test.
It will not be a test. It will be whatever the model happened to surface, shaped by whoever was standing there, and nobody in the room will know what is missing.
The years are the thing being replaced, and the years are the part that matters. An engineer who has run hundreds of engagements knows which attack path is worth pulling on because they have seen it fail before. They recognize the thing that is only suspicious if you know how that technology is normally deployed. They notice what is absent. None of that is in the model — it is what tells the operator where to point it.
So a company saves money, gets a clean-looking report, and stays exposed on the path nobody thought to look at. They will not find out until somebody else does.
Until something genuinely replaces an engineer — and nothing does yet — human-led testing has to be the standard. This credential is how you tell the difference between an engineer amplified by AI and somebody being handed a result.
The other half of it
Used properly, they miss far less than they would alone.
An engineer working by hand has to choose what to look at. There is not enough time in an engagement to read every code path, test every parameter on every endpoint, and chase every variant of a root cause across infrastructure — so they prioritize, and prioritizing means deciding in advance what to ignore.
Used properly, AI removes most of that compromise. The same engineer can cover volume no human could work through unaided — every request, every parameter, every reachable path — while still applying the judgment that decides what actually matters and what a finding really means.
That is the whole point of the credential. Not faster reports. The same expert judgment, applied across a far larger surface, with the client’s data staying where it belongs while it happens.
Better, faster, and safer for your data — because the engineer decides what matters and the model covers ground they could never cover alone.
The missing link
The problem is not that engineers use AI. It is that most use it badly.
The public argument is still stuck on whether generative AI belongs in security testing. That argument is a distraction, because the people causing the damage are not waiting for it to be settled. They are already using it, with no structure around it, and the results are going out to clients.
Point a model at an engagement with nothing enforcing rigor and it will produce volume — plausible findings, confident severity ratings, remediation advice, and a substantial amount that is simply not true. An engineer who cannot tell the difference passes it straight through. The client gets a longer report and a worse test.
That is where the false positives come from, and it is why so many buyers have concluded that AI-assisted testing is noise. They are drawing a fair conclusion from real evidence. They have just never been shown the other version.
Because used correctly, it is the largest coverage gain this work has ever had, with an experienced engineer deciding what any of it means. Used incorrectly, it is a false-positive generator with excellent grammar. Nothing on the outside of a report tells you which one you paid for.
That is the gap this credential closes. Not whether they used AI — whether they are any good at it.
Where this comes from
We have been testing with models since before they were any good at it.
Security Arsenal has been running client engagements with generative AI in the loop for three years — since roughly the point it became possible at all. In most fields three years is nothing. In this one it is the entire history of the thing.
It did not start well. Early on it was like teaching a toddler to code: you would get halfway through explaining the problem and a squirrel would go by. Everything we now know about keeping a model on task, bounded to its scope, and honest about what it actually verified was learned by doing it badly first — on our own engagements, and fixing it.
That is where this exam comes from. Not a curriculum somebody wrote about AI, and not a vendor course. Three years of finding out which parts of this genuinely work, which parts quietly produce garbage, and what has to be true of an engineer before either one is safe in front of a client.
For whoever is paying
What a client gets to conclude.
Most of this page is written for the engineer sitting the exam. If you are the one reading APEX on a proposal and deciding whether it means anything, these are the claims — each tied to the specific thing the candidate had to survive.
The engineer is doing the work, not the model.
Unplug the AI and they can still run your engagement. It takes longer, and more of your infrastructure goes unread — which is the honest reason the good ones use it.
They get through your whole infrastructure, not a sample of it.
Without AI, somebody decides what to skip and you never hear which parts. With it used properly, far less gets skipped — and the report tells you what did.
They understand your infrastructure, not just a methodology.
You cannot break what you cannot build. They built it.
Your data does not end up in somebody’s chatbot.
There is finally a credential that fails people for the thing a lot of testers are quietly doing right now.
They find what a scan never will.
We do not take the easy way into your network. A CVE scan can only ever find what somebody else already published.
Prove it — every finding has a working exploit behind it.
Put your skills behind your words. And if an APEX holder could not exploit it, very few people or systems were going to.
The report tells you what it did not test.
A clean report that quietly skipped half your infrastructure is the thing that gets you breached.
The full version, written for clients — including the three different things that get sold as a penetration test, and six questions worth asking any tester, ours or anybody else’s.
Before you go any further
This is not a penetration-testing course. If you cannot run the engagement, write the tooling, prove the exploit and defend the report with every model switched off, you are not ready for this yet.
We would rather cost you five minutes on a free page than a failed exam. There is a self-check that runs entirely in your browser and is perfectly willing to tell you no.
The structure
You can’t break what you can’t build.
Step one
You build the infrastructure.
Cloud and on-premise virtualization, a routed core, remote access, perimeter, a mixed operating-system fleet, managed mobile devices and a working voice platform. You design it, you configure it, you defend the design.
Step two
You build the tooling.
A clean Kali server and one AI API token. No agent framework, no prepared prompts, no Security Arsenal software. You build the thing that does the testing — and the controls that stop client data reaching the model while it works.
Step three
You attack a different one.
Never your own build. You draw an anonymized infrastructure from the graded pool and go at infrastructure a real engineer configured — with all the inconsistency that implies.
What you have to be able to stand up
Cloud and virtualization
- AWS
- Azure
- VMware
- Proxmox
Network and edge
- BGP
- OSPF
- SDN
- VPN
- Firewall
- Cloudflare
- CDN
Endpoints and servers
- Windows
- Linux
- macOS
- Endpoint security
- Security monitoring
Mobile and voice
- Apple mobile
- Android
- VoIP
Every component is graded on whether it works and whether you can explain why you built it that way — not on how hard it is to break. Details on the build-and-break page.
What it is not
The exclusions are the product.
- Excluded:Not beginner penetration-testing instruction, and not a "become a hacker" program.
- Excluded:Not a certification for running a scanner, pasting the output into a model, or matching software versions to CVEs.
- Excluded:No pass based on multiple-choice questions, attendance, course completion, or an AI-generated report.
- Excluded:No uncontrolled use of consumer or public AI services with client code, credentials, screenshots, logs or findings.
- Excluded:No finding accepted without a reproducible attack path, evidence, impact analysis and source or runtime reasoning.
This is not a credential about attacking AI systems.
There are good credentials that certify testing AI systems themselves — model behavior, prompt handling, agent abuse, the AI attack surface. APEX measures the other direction: an engineer using AI to test everything else, on real client work, without losing confidentiality, technical accuracy or personal accountability. The two do not overlap much, and holding one says nothing about the other.
The performance exam
7 phases. The models are off for two of them.
Phase A: Manual foundation
AI disabledDo the work with the models switched off.
Deliverable Manual exploitation, code modification, attack-path notes, live spot checks.
Phase B: Build the infrastructure
Allowed through the controlled gatewayDesign, build and configure a full enterprise infrastructure from bare metal up.
Deliverable A running infrastructure, its configuration as code, a design record, and a build conformance run that passes.
Phase C: Defend what you built
Allowed through the controlled gatewayThe swarm attacks your infrastructure. You are sitting in front of it.
Deliverable A live defense of your own build, your detection and response record, and an account of what got through and why.
Phase D: Build the tooling
Approved sandbox onlyA bare Kali server and an AI API token. Nothing else.
Deliverable A candidate-built AI security application, its threat model, its policy enforcement and its test suite.
Phase E: Attack an environment you have never seen
Allowed through the controlled gatewayUnique to you, built to be unfamiliar, and never used twice.
Deliverable Original findings, proofs of concept, an evidence graph and a coverage ledger.
Phase F: Reporting
Allowed, fully loggedEvidence-locked reports, plus a full account of what you did with AI.
Deliverable Executive and technical reports, an evidence package, a coverage ledger, an AI-use manifest and a validation manifest.
Phase G: Oral defense
AI disabledExplain it live, with the models off.
Deliverable Live explanation, replay, challenge questions, remediation defense.
Exam length is being set from pilot data rather than from a marketing page. When we have run enough of them to know, we will publish the number. Full exam structure.
Recognized prior work
Hold an OSEE? One phase comes off. The rest does not.
Waived
Phase A — Manual foundation
You have already demonstrated it, under exam conditions, to a standard we respect.
Still required
- Building the infrastructure
- Voice
- Software development
- Every AI phase
Security Arsenal recognizes OSEE as evidence for our manual foundation gate. This is our own admissions decision. It is not a partnership, endorsement, affiliation or articulation agreement with OffSec, and it confers nothing on their behalf. How the bridge route works.
Client-data protection
A penetration test contains exactly what an attacker wants.
| Class | Examples | Default AI destination |
|---|---|---|
| Public | Published domains, public source, public documentation. | Approved enterprise or private model. A public model only where company policy explicitly allows it. |
| Internal | Generic methods, sanitized templates, non-client operational notes. | Approved enterprise tenant with contractual controls. |
| Client confidential | Architecture, code, findings, screenshots, logs, asset names. | Security Arsenal-controlled or client-approved private deployment. No consumer or public endpoint. |
| Restricted | Credentials, tokens, PII and PHI, private keys, production data, regulated records. | Prefer no model ingestion at all. If essential: isolated approved deployment, minimized, encrypted, audited, explicitly authorized. |
| ProhibitedBlocked | Data barred by contract or law, another client’s data, unapproved secrets, exam answer material. | Never submitted. Blocked and alerted. |
Candidates build and pass six leakage tests, including cross-tenant retrieval, canary strings and prompt-injection exfiltration. The full control standard.
Low false positives by design
A fluent explanation is not a finding.
| State | Definition | In the report? |
|---|---|---|
| Hypothesis | A potential issue proposed by a human, a model, an analyzer or a test. | Not in the report |
| Candidate | Code or runtime evidence suggests a plausible attack path. | Not in the report |
| Validated | Independent reproduction crosses the claimed security boundary. | In the report |
| Conditionally validated | Exploitability is proven under clearly stated conditions; constraints remain. | In the report, conditions stated up front |
| Rejected | Negative testing disproved the claim, or the evidence was insufficient. | Not in the report; retained internally for scoring |
| Duplicate or variant | Same root cause or attack primitive as another finding. | Merged; affected paths preserved |
Key doctrine
AI confidence is not evidence. Multiple agents repeating the same unsupported claim are not independent confirmation. A finding exists only when the candidate can show the affected asset and code path, the preconditions, the controlled action, the observed result, the security boundary crossed, the impact, and a repeatable method that reproduces it.
The easy way in
We document the CVE. We write our own exploit.
A CVE scan tells you which of your software versions appear in a public list. That is inventory work, it is genuinely useful, and it is not a penetration test. Run as one, it can only ever surface what somebody else already found and published — so it will never once find the thing being used against you by an attacker who did their own work.
Where a known vulnerability is present, it gets documented, because you need to know. But the exploitation is original. A published proof of concept establishes that a version matched. Writing our own establishes that it is genuinely reachable in your environment, past your controls, in your configuration — which is the only version of that question you actually care about.
The harder reason is what it does to the engineer. Somebody who only ever runs other people’s exploits never builds the ability to find the one nobody has published. That ability is the entire job, and it is exactly what the exam is set up to measure.
| Qualifies | Does not |
|---|---|
| A business-logic or authorization flaw derived from code and runtime behavior. | A scanner saying IDOR or BOLA may exist. |
| An unsafe interaction between components that are individually reasonable. | A dependency with a known CVE. |
| A state-machine, race, cache, tenancy, trust-boundary, parser or workflow weakness. | A generic missing-header observation. |
| A candidate-built proof crossing a security boundary with controlled impact. | Model-generated exploit text that was never understood or executed. |
| Variant analysis showing related reachable paths and a shared root cause. | One copied request with no explanation. |
This is why the exam requires an original, non-public discovery derived from code, architecture, state or runtime behavior — and why a scanner result, a version match or a copied exploit is explicitly not accepted as a finding.
The journey
Course completion is not certification.
1Apply
Submit experience, languages, build and engagement history, and identity.
Outcome Eligibility review. Nothing is sold before prerequisites are clear.
2Baseline challenge
Complete a short manual coding and security reasoning screen.
Outcome Admit, require specific remediation, or decline.
3Security AI Engineering
Work through privacy architecture, model and tool engineering, evidence discipline, orchestration and reporting.
Outcome Course completion record only — not a certification.
4Practice range
Build and harden AI testing applications, and complete replayable scenarios.
Outcome Readiness signals and an evidence portfolio.
5Exam readiness
Meet the objective readiness gates and self-attest.
Outcome Performance exam scheduled.
6APEX exam
Build, tool, attack, report and defend.
Outcome Pass, fail, or a limited retest decision.
7Credential
Accept the code of conduct and set publication preferences.
Outcome Verifiable credential issued.
8Internal authorization
Security Arsenal employees only: complete company controls and supervised operation.
Outcome Revocable SwarmPT Lead Authorization.
Who it is for
We built this for the whole profession. Including our competitors.
When an AI-assisted test is inaccurate or incomplete, the client carries the risk. Not the firm that wrote the report. Not the engineer who ran it.
That is a bad arrangement, and it does not get better if the standard of care only exists inside one company. So APEX is open to any qualified practitioner, whoever they work for — including firms that compete with us directly, for the same clients, in the same week.
Grading is anonymized. A reviewer sees a candidate handle, never a name or an employer, so a decision about a competitor’s engineer cannot be questioned and cannot be leaned on.
We would rather compete against engineers who are held to this standard than win against ones who are not.
Working practitioners
Testers, red teamers, application security engineers, exploit developers and senior security engineers who already run authorized engagements.
Internal security teams
People who test their own infrastructure and need a defensible standard for what may go into a model and how a finding is established.
Testing firms
Including ours. Our engineers sit the same exam, graded the same way, against the same rubric, with the same handle instead of a name.
The price, and what is still open
The price is published. So is the list of things we still do not know.
Next step
Find out whether this is for you — before you give us anything.
The eligibility check runs entirely in your browser. No email box, no gated result, no follow-up sequence. It is capable of telling you that you are not the intended level, and it will.
Straight answers
The questions people actually ask.
- What is APEX?
- APEX — the Advanced Penetration Engineering Expert, awarded by Security Arsenal — is a performance-based certification for experienced penetration engineers who use AI on authorized client work. It measures manual capability with AI switched off, infrastructure engineering, secure AI architecture, original code-derived vulnerability discovery, evidence-based reporting and a live oral defense.
- Is APEX a certification about hacking AI systems?
- No. There are separate credentials that certify testing AI systems themselves — model behavior, prompt handling, agent abuse. APEX measures the other direction: an engineer using AI to test everything else, without losing confidentiality, technical accuracy or personal accountability. Holding one says nothing about the other.
- Is APEX suitable for beginners?
- No. It assumes the candidate can already conduct a complete engagement, develop tooling, prove an exploit and defend a report without any generative AI. A candidate who cannot do that does not pass the manual foundation gate, which no score elsewhere compensates for.
- Why does the APEX exam make you build a network before attacking one?
- Because you cannot break what you cannot build. The candidate designs and configures a full enterprise infrastructure — cloud and on-premise virtualization, BGP and OSPF routing, VPN, firewall, Windows, Linux and macOS fleets, managed Apple and Android devices, and a working VoIP platform — and is then assigned a different candidate’s infrastructure to attack. Building it proves the understanding that tells an engineer where to look. The harder lesson is how infrastructure actually gets misconfigured. Most real environments were stood up by somebody who got them working under time pressure without fully knowing how to secure them, and working and secure are two different jobs. A candidate only learns to recognize that by building the thing themselves — enabling a setting because something would not work otherwise, then attacking their own design to discover what that setting just handed away. One convenience option, turned on by an administrator who needed the ticket closed, is regularly the entire engagement. Having sat on both sides of the same infrastructure, an APEX holder can both detect that decision and advise how to close it without breaking what it was doing.
- Does an OSEE waive part of the APEX exam?
- An active OSEE (OffSec EXP-401, Advanced Windows Exploitation) satisfies the manual foundation gate and waives that phase only, once the certification number is verified with the issuing body. It does not waive the build phase, the VoIP requirement, the software development requirement or any AI phase. This is Security Arsenal’s own admissions decision, not a partnership or endorsement.
- Can engineers at competing security firms earn APEX?
- Yes, deliberately. When an AI-assisted test is inaccurate or incomplete the client carries the risk, and that does not improve if the standard of care exists inside only one company. Grading is anonymized — a reviewer sees a candidate handle, never a name or an employer — so a decision about a competitor’s engineer cannot be questioned and cannot be leaned on. Security Arsenal engineers sit the same exam, graded the same way, against the same rubric.
- Does completing the Security AI Engineering course make you APEX certified?
- No. Course completion produces a separate record with its own identifier and expiry, marked as not being the certification. Certification comes only from passing the independent performance exam, and many people who complete the course will not pass it.
- Why does another security certification need to exist?
- Existing certifications measure the job as it was before a model sat next to the engineer. Two questions now decide whether a test was any good, and none of them ask either: can this person still do the work with AI switched off, and where did the client’s data go? A penetration test contains source code, credentials, architecture and a list of exactly how to get in. Any existing certification can be held in good standing by someone pasting that into a consumer chatbot, and none of them would notice.
- What does it mean for a client if their penetration tester holds APEX?
- Four things. The engineer can run the engagement with generative AI switched off — one exam phase and the oral defense are conducted that way, as a gate no other score compensates for. They can build the infrastructure they are attacking, because the exam makes them do it first. Client data stays inside approved systems, because they built and then defended those controls under adversarial test, and a confirmed leak ends the attempt. And every finding survived independent reproduction, with the report stating explicitly what was not tested.
- Can AI replace a penetration tester?
- Not yet, and that is the problem APEX exists to address. AI tooling makes it possible to hand a penetration test to somebody who has never run one — point the right setup at a network and findings, severities and a report come out. That is a result, not a test. An engineer with years of engagements knows which attack path is worth pulling on, recognizes what is only suspicious if you know how the technology is normally deployed, and notices what is absent — none of which is in the model. The reverse is also true and less often said: an engineer working without AI misses things too, because an engagement produces more data than a person can read, so they cover a subset and the rest goes untested. Used properly, AI removes most of that limit — the same engineer covers every request, every parameter and every reachable path while still supplying the judgment about what any of it means. Neither half works alone. Until something genuinely replaces an engineer, human-led testing has to be the standard.
- Why do AI-assisted penetration tests produce so many false positives?
- Because most of the people using AI on engagements are using it badly, not because the approach does not work. Point a model at a target with nothing enforcing rigor and it produces volume — plausible findings, confident severity ratings and a substantial amount that is simply untrue. An engineer who cannot evaluate that output passes it straight to the client, who receives a longer report and a worse test. Buyers who conclude from this that AI-assisted testing is noise are drawing a fair conclusion from real evidence; they have just never been shown the version where it is done correctly. APEX exists to make that difference checkable: the exam requires a working proof of concept for every finding, independently reproduced by a reviewer, and a fluent explanation with no artifact behind it is scored as a rejection.
- Is it safe for a penetration tester to use AI on my data?
- It depends entirely on where the data goes, not on whether a model was involved. A private, self-hosted or contractually controlled deployment raises none of the usual concerns. A consumer chatbot raises all of them — a penetration test contains source code, credentials, architecture, screenshots, logs and a list of exactly how to get in, and that material is worth the same to an attacker wherever it ends up. Every existing security certification can be held in good standing by someone pasting client data into a public model today, and none of them would notice. APEX candidates build the controls that enforce the distinction — classification, outbound inspection, per-engagement isolation, canaries, fail-closed routing — and then those controls are attacked. A confirmed leak ends the exam attempt outright.
- How long does an APEX certification last, and how do you renew it?
- Two years. Three was too long — the models change, their failure modes change and the defaults on the infrastructure being tested change, so a credential issued against this year's landscape cannot honestly describe a much later one. Renewal has two halves and is deliberately not a repeat of the exam: evidence of 40 hours of real work across the term — engagements delivered, original research, tooling other people ran, teaching, with structured training capped because attendance is the weakest evidence — plus a half-day delta assessment covering only what changed since the last sitting. Renewal is $950 against a $6,500 exam, because a renewal priced like a second exam is one nobody takes, and a lapsed credential proves nothing about anybody.
- How do you verify an APEX credential?
- At securityarsenal.com/verify, using both the holder’s name and their credential number. Both are required: a credential number on its own returns nothing, so the registry confirms what a verifier already knows rather than acting as a browsable directory of certified engineers. Verification shows the current status — active, expired, suspended or revoked — never exam scores or employer.
- What does the APEX exam cost and how long does it take?
- The exam is $6,500, or $5,800 on the OSEE bridge route. The Security AI Engineering course is $2,950, and the course-plus-exam bundle is $8,500. Eligibility review is free, permanently — nobody pays to be told they are not ready. Retakes start at $300 because a retake priced like a second exam just loses the candidate. Exam length is not published yet: it is being set from pilot evidence rather than from a marketing page. Full pricing is at securityarsenal.com/certifications/sa-apex/pricing.
- Does APEX require working exploit code for every finding?
- Yes. Working code, or it is not a finding. Every issue submitted in the exam carries a proof of concept the candidate wrote, that a reviewer executes independently in a clean environment, and that demonstrably crosses the security boundary being claimed. A description of an attack, a scanner result with a CVSS score, a screenshot, a request copied from a write-up, or model-generated exploit code the candidate cannot explain are all rejected. The oral defense then samples that code line by line with generative AI disabled. The inverse also carries weight: when a holder could NOT exploit something, that is meaningful, because the same exam established they can direct AI across an entire attack surface.
- What ends an APEX attempt immediately?
- Using AI to explain code the candidate cannot independently explain; uploading protected data to an unapproved endpoint or a public model; acting outside scope or attacking shared infrastructure; claiming scanner or model output as an exploitable finding without proof; fabricating evidence; missing a designated critical issue that the coverage map makes reasonably discoverable; or producing a materially unsafe exploit or remediation.