For years, the defensive community has debated when — not if — frontier AI models would cross the line from assisting with vulnerability research to autonomously producing working exploits. Anthropic's Frontier Red Team just gave us a data point that should end the debate.
In a recently published evaluation, Anthropic ran several models against 100 randomly selected tasks from an internal binary security issue benchmark. The results: GLM-5.3 developed full control flow hijacks in 4% of trials, and Claude Mythos Preview did so in 6%. The critical sentence in the report is the follow-on: earlier models — Claude Opus 4.6 and GLM-5.2 — succeeded in none of them.
Let me translate that into operational terms. A control flow hijack is the pivotal moment in binary exploitation where an attacker redirects execution — the foundation for arbitrary code execution. A generation ago, these models couldn't do it at all. Now, two independently developed models from different vendors do it reliably enough to show up in single-digit percentages across a randomized benchmark. That trajectory is what matters. Single-digit success rates at scale, with unlimited retries and parallelized attempts, is not a curiosity — it's a capability.
Why Defenders Should Treat This as an Operational Warning, Not a Research Curiosity
I've led IR engagements where the exploit used against the client was clearly hand-crafted by a skilled human — weeks of work, obvious expertise. The economic reality of offensive operations has always been a rate limiter: skilled exploit developers are rare, expensive, and slow. That rate limiter is now eroding.
Here's what this specific finding means for your defensive posture:
1. The window between vulnerability disclosure and working exploit is compressing. If a model can autonomously develop control flow hijacks against binary targets, the labor bottleneck in exploit development shrinks. Defenders who historically had days-to-weeks of breathing room after a patch drops must now assume hours-to-days — and that assumption should already be driving your patch prioritization SLAs.
2. Capability diffusion is the real story. Anthropic's framing — "the spread of advanced cyber capabilities" — is deliberate. This isn't about one vendor's model. GLM-5.3 is a non-US-developed model, and it crossed the threshold too. You cannot regulate, geofence, or terms-of-service your way out of this. Multiple model families now possess the capability, and it will propagate to open-weight and fine-tuned derivatives.
3. A 4-6% success rate is meaningless at machine scale. Humans get tired. Models don't. An operator (or eventually an autonomous agent) running thousands of parallel attempts against a target binary will reach a working exploit. Probabilistic success at low per-trial rates is still deterministic success at volume.
4. Attack volume against memory-corruption classes will rise. Buffer overflows, use-after-free, and type confusion bugs in network-facing services, browsers, and parsers have always required elite skill to weaponize. Expect more exploitation attempts against these classes — and expect more of those attempts to be noisy, iterative, and crash-heavy during the AI's development loop. That noise is your detection opportunity.
Technical Analysis: What a Control Flow Hijack Capability Implies for Your Telemetry
To be clear about scope: Anthropic's benchmark tasks are controlled binary exploitation exercises — CTF-adjacent challenges, not confirmed in-the-wild exploitation of production software. There is no CVE, no active exploitation campaign, and no CISA KEV entry tied to this announcement. This is a capability inflection point, not an active incident.
But capability inflection points are exactly when defenders should retool. Here's the exploitation mechanics from a defender's lens:
How the attack chain works at the endpoint level:
- Input delivery — a crafted input (network request, file, message) reaches a vulnerable parser or service.
- Memory corruption — the input triggers a bug (overflow, UAF, OOB write) that corrupts program state.
- Control flow hijack — corrupted state (a return address, function pointer, or vtable) redirects execution. This is the step the models now autonomously achieve.
- Payload execution — shellcode or a ROP chain runs, typically spawning a child process, loading a module, or establishing a C2 channel.
The critical detection insight: AI-driven exploit development is iterative. A model developing an exploit against your exposed service will cause repeated crashes before it succeeds — and even a successful exploit often crashes the host process first. Automated exploitation at scale produces an observable signature: crash loops, anomalous process restarts, Windows Error Reporting artifacts, crash dumps, and protected services spawning unexpected children. Human exploit developers test quietly in a lab; an automated pipeline hammering your internet-facing service is far noisier. Instrument for the noise.
Affected surface to prioritize: internet-facing network services, VPN/remote access gateways, email and file-parsing infrastructure, browser fleets, and any custom or legacy binary services that haven't seen a security review in years. Legacy C/C++ services with no modern mitigations (DEP, ASLR, CFG) are the softest targets for an AI still climbing the exploitation learning curve.
Detection & Response
The rules below target the observable artifacts of exploitation attempts against binary services — crash telemetry, crash-followed-by-execution patterns, and exploit mitigation events. These are tuned to fire on the noisy, iterative behavior that automated exploitation produces, not on generic admin activity.
Sigma Rules
---
title: Network Service Crash Followed by Child Process Execution
title_note: Potential successful control flow hijack
id: 8f2a1c44-3b7e-4d91-a2c6-9e0f5d6b7a31
status: experimental
description: Detects a network-facing service process spawning an unexpected child process shortly after service instability, consistent with successful memory corruption exploitation and control flow hijack.
references:
- https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
- https://attack.mitre.org/techniques/T1190/
- https://attack.mitre.org/techniques/T1059/
author: Security Arsenal
date: 2026/09/30
tags:
- attack.initial_access
- attack.t1190
- attack.execution
- attack.t1059
logsource:
category: process_creation
product: windows
detection:
selection_parent:
ParentImage|endswith:
- '\w3wp.exe'
- '\httpd.exe'
- '\nginx.exe'
- '\svchost.exe'
- '\sqlservr.exe'
- '\lsass.exe'
- '\spoolsv.exe'
selection_child:
Image|endswith:
- '\cmd.exe'
- '\powershell.exe'
- '\pwsh.exe'
- '\wscript.exe'
- '\cscript.exe'
- '\mshta.exe'
- '\rundll32.exe'
- '\regsvr32.exe'
- '\certutil.exe'
- '\bitsadmin.exe'
condition: selection_parent and selection_child
falsepositives:
- IIS application pools legitimately spawning management tooling (rare; baseline per host)
- SQL Server executing xp_cmdshell (itself a high-risk finding worth investigating)
level: critical
---
title: Repeated Application Crash Loop Indicating Automated Exploit Development
id: 3c9d7e21-5a48-4f0b-b1d3-7c2e8a9f4b56
status: experimental
description: Detects repeated crashes of the same application via Windows Error Reporting, a hallmark of iterative, automated exploit development against a target binary. Tune threshold logic in your SIEM to group by executable within a short window.
references:
- https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
- https://attack.mitre.org/techniques/T1499.004/
author: Security Arsenal
date: 2026/09/30
tags:
- attack.impact
- attack.t1499.004
logsource:
product: windows
service: application-error
detection:
selection:
EventID: 1000
filter_known_noisy:
Image|endswith:
- '\chrome.exe'
- '\msedge.exe'
- '\firefox.exe'
condition: selection and not filter_known_noisy
falsepositives:
- Unstable legitimate software; the signal is repetition of crashes on a server-class binary, not a single crash
level: medium
---
title: Exploit Guard Exploit Protection Mitigation Triggered
id: b14e6f08-2d9c-4a75-83e1-6f0a3c9d5e27
status: experimental
description: Detects Windows Defender Exploit Guard blocking an exploitation attempt (ACG, ASLR, DEP, control flow violations) against a protected process — direct evidence of control flow hijack attempts against hardened targets.
references:
- https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
- https://attack.mitre.org/techniques/T1211/
author: Security Arsenal
date: 2026/09/30
tags:
- attack.defense_evasion
- attack.t1211
logsource:
product: windows
service: security-mitigations
detection:
selection:
EventID:
- 1 # ACG audit/block
- 2 # ASLR violation
- 3 # DEP violation
- 11 # Control flow violation
condition: selection
falsepositives:
- Legacy applications incompatible with enforced mitigations during initial Exploit Guard rollout (run in audit mode first)
level: high
KQL — Microsoft Sentinel / Defender Hunt
This query hunts the specific exploitation pattern: a server-class process crashing and then being observed spawning command interpreters or loading anomalous modules — the fingerprint of a control flow hijack that succeeded after iterative attempts. It also surfaces crash-loop behavior that indicates an automated exploitation pipeline is actively working your perimeter.
// Hunt: crash loops and post-crash execution on server-class processes
// Indicates iterative/automated exploit development or successful control flow hijack
let ServerProcesses = dynamic(["w3wp.exe","httpd.exe","nginx.exe","sqlservr.exe","spoolsv.exe","svchost.exe","node.exe","java.exe","tomcat9.exe"]);
let SuspiciousChildren = dynamic(["cmd.exe","powershell.exe","pwsh.exe","mshta.exe","wscript.exe","cscript.exe","rundll32.exe","regsvr32.exe","certutil.exe"]);
// Part 1: Crash loop detection - same process crashing repeatedly in 1h window
let CrashLoops = DeviceEvents
| where TimeGenerated > ago(24h)
| where ActionType in ("AppCrash", "ProcessCrash") or (FileName in~ ServerProcesses and ActionType contains "Crash")
| summarize CrashCount = count(), FirstCrash = min(TimeGenerated), LastCrash = max(TimeGenerated) by DeviceName, FileName, bin(TimeGenerated, 1h)
| where CrashCount >= 3;
// Part 2: Server process spawning suspicious child (post-exploitation execution)
let PostExploitExec = DeviceProcessEvents
| where TimeGenerated > ago(24h)
| where InitiatingProcessFileName in~ ServerProcesses
| where FileName in~ SuspiciousChildren
| project DeviceName, InitiatingProcessFileName, InitiatingProcessCommandLine, FileName, ProcessCommandLine, AccountName, TimeGenerated;
union CrashLoops, PostExploitExec
| order by DeviceName, TimeGenerated asc
Velociraptor VQL — Endpoint Crash Artifact Hunt
When you suspect an automated exploitation pipeline has been working a host, pull the crash dump and WER artifacts directly — they tell you which binary was targeted, how many attempts occurred, and preserve the corrupted memory state for forensic analysis of the exploit itself.
-- Hunt for Windows Error Reporting crash artifacts indicating exploitation attempts
-- against server-class binaries (iterative crash loops from automated exploit dev)
SELECT FullPath, Mtime, Size, basename(path=FullPath) AS DumpFile
FROM glob(globs='C:/ProgramData/Microsoft/Windows/WER/**/Report*.wer')
WHERE Mtime > now() - 86400 * 7
UNION
SELECT FullPath, Mtime, Size, basename(path=FullPath) AS DumpFile
FROM glob(globs='C:/Users/*/AppData/Local/CrashDumps/*.dmp')
WHERE Mtime > now() - 86400 * 7
UNION
-- Correlate with live network listeners on server processes spawning shells
SELECT Pid, Name, CommandLine, Exe, Username, CreateTime
FROM pslist()
WHERE (Name =~ '(?i)cmd|powershell|pwsh|mshta|rundll32')
AND CreateTime > now() - 86400
Hardening Script — Enable Exploit Mitigations on Critical Services
The single highest-value defensive action against autonomous exploit development: make exploitation expensive. Modern mitigations (DEP, ASLR, CFG, ACG) force the attacking model to solve a much harder problem, burn more attempts, and generate more crash telemetry for your SOC.
#Requires -RunAsAdministrator
# Harden internet-facing and legacy binaries with Exploit Protection mitigations
# Run in audit mode first, review Microsoft-Windows-Security-Mitigations logs, then enforce
$targets = @(
"C:\inetpub\w3wp.exe",
"C:\Windows\System32\spoolsv.exe",
"C:\Program Files\YourLegacyApp\service.exe" # Replace with actual legacy binaries
)
foreach ($exe in $targets) {
if (Test-Path $exe) {
Write-Host "[+] Applying mitigations to $exe" -ForegroundColor Cyan
# DEP (permanent), ASLR (force relocate, bottom-up), CFG, SEHOP, ACG
Set-ProcessMitigation -Name (Split-Path $exe -Leaf) -Enable DEP, EmulateAtlThunks, BottomUp, ForceRelocateImages, HighEntropy, SEHOP, CFG, StrictCFG, BlockDynamicCode, AllowThreadOptOut, AllowRemoteDowngrade
} else {
Write-Warning "[-] Not found, skipping: $exe"
}
}
# Verify applied mitigations
foreach ($exe in $targets) {
if (Test-Path $exe) {
Write-Host "`n[=] Mitigations for $(Split-Path $exe -Leaf):" -ForegroundColor Green
Get-ProcessMitigation -Name (Split-Path $exe -Leaf) | Format-List
}
}
# Ensure WER crash reporting is capturing artifacts for forensic review
$werPath = "HKLM:\SOFTWARE\Microsoft\Windows\Windows Error Reporting\LocalDumps"
if (-not (Test-Path $werPath)) { New-Item -Path $werPath -Force | Out-Null }
Set-ItemProperty -Path $werPath -Name "DumpFolder" -Value "C:\CrashDumps"
Set-ItemProperty -Path $werPath -Name "DumpCount" -Value 10 -Type DWord
Set-ItemProperty -Path $werPath -Name "DumpType" -Value 2 -Type DWord # 2 = Full dump
New-Item -Path "C:\CrashDumps" -ItemType Directory -Force | Out-Null
Write-Host "[+] WER full dumps configured at C:\CrashDumps (rotate and restrict ACLs)" -ForegroundColor Cyan
Remediation & Strategic Hardening
There is no patch for an adversary capability — but there is a concrete defensive program this announcement should trigger:
-
Compress your patch SLAs now. If your organization still patches internet-facing criticals on a 14-30 day cycle, that window no longer matches adversary economics. Target 72 hours for internet-facing criticals, 24-48 hours for anything with a public PoC. Re-baseline your vulnerability management program against this assumption.
-
Enforce exploit mitigations on every binary you can't patch. Legacy services that can't be updated quickly become the preferred target for an adversary whose marginal cost per attempt is near zero. DEP, ASLR, CFG, and ACG materially raise the difficulty bar (script above).
-
Instrument crash telemetry as a first-class detection source. Most SOCs treat application crashes as an IT problem. In an era of automated exploitation, crash loops on server binaries are an intrusion signal. Route WER, core dumps, and service restart events into your SIEM with correlation logic.
-
Assume iterative attack noise and alert on it. Automated exploit development retries aggressively. Rate-based detections on service crashes, malformed request floods against specific parsers, and repeated segmentation faults from the same source are high-fidelity in a way they weren't three years ago.
-
Reduce your binary attack surface. Every internet-facing custom binary, abandoned plugin, and unreviewed parser is target practice for a system that never gets tired. Decommission, isolate, or put strict egress controls around anything that can't be hardened.
-
Update your threat model and tabletop scenarios. Run an exercise where the adversary generates a working exploit against your stack within 48 hours of a disclosure. Most organizations' response plans implicitly assume more time than that. Find the gaps before someone else's pipeline does.
-
Track frontier capability evaluations as threat intelligence. Anthropic, and increasingly other labs, publish these red team assessments. Treat capability evaluations the way you treat KEV additions — as triggers to revalidate your exposure and detection coverage.
The models crossed the threshold at single-digit success rates this quarter. The defensive community's job is to make sure that by the time those rates hit double digits, the noise those attempts generate is already lighting up our consoles.
Related Resources
Security Arsenal Red Team Services AlertMonitor Platform Book a SOC Assessment pen-testing Intel Hub
Is your security operations ready?
Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.