A quiet shift has occurred in enterprise endpoints over the last 18 months: developers and engineers are running autonomous and semi-autonomous AI coding agents — Claude Code, OpenAI Codex CLI, Gemini CLI, Cursor, GitHub Copilot, Warp, Windsurf, Qwen Code, OpenCode, and emerging tools like Hermes — directly on workstations that hold source code, credentials, and production access. These agents don't just suggest code. They read files, write files, execute shell commands, spawn subprocesses, and in some configurations reach out to MCP servers and external APIs with the user's privileges.
SANS recently updated FOR577 (Linux/Windows forensic analysis for incident response) with new day-5 material covering investigation of AI usage in IR, walking through the eight most popular coding assistants and where their evidence lives on disk. The accompanying research into OpenCode chat history and Hermes artifact locations underscores the core problem: every one of these agents leaves a forensic trail — chat transcripts, session logs, tool-call records — but almost no SOC is collecting it.
Why defenders should care right now:
- Prompt injection is an execution vector. Malicious instructions embedded in repos, READMEs, web pages, or MCP server responses can cause an agent to exfiltrate data, modify code, or run commands — and the resulting activity looks like legitimate user behavior.
- Shadow AI is an inventory problem. Agents installed outside IT approval bypass your EDR policy assumptions, your DLP rules, and your software allowlists.
- Post-incident reconstruction is failing. When a breach involves a developer workstation, investigators increasingly find agent-executed commands in shell history with no corresponding human session. If you can't tie an action to an agent session, you can't establish intent or scope.
This post maps the artifact locations, gives you deployable detections for agent-driven execution, and provides an audit script to inventory AI agents across your fleet.
Technical Analysis
The Threat Model: Agents as Privileged Executors
AI coding agents share a common architecture that matters forensically:
- A local CLI or IDE-integrated process (e.g.,
claude,codex,gemini,opencode, or Electron apps like Cursor and Windsurf). - A tool-call execution layer that can spawn shells (
cmd.exe,powershell.exe,/bin/bash,/bin/zsh), run interpreters (node,python), and read/write arbitrary files within user context. - A persistent session/transcript store on disk — typically JSONL or SQLite — recording prompts, tool calls, file diffs, and command outputs.
- Optional MCP (Model Context Protocol) integrations that extend agent reach to local services, databases, and SaaS APIs.
From a defender's perspective, the critical observation is that agent-executed commands have a distinct process lineage: the agent process is the parent. That lineage is your highest-fidelity detection signal, and the on-disk transcript is your reconstruction source.
Key Forensic Artifact Locations
Based on published research and hands-on validation, the following locations hold session history and evidence:
| Agent | Primary Artifact Locations |
|---|---|
| Claude Code | %USERPROFILE%\.claude\projects\<project-slug>\*.jsonl (Windows), ~/.claude/projects/ (macOS/Linux), ~/.claude.json config, ~/.claude/history |
| OpenAI Codex CLI | ~/.codex/sessions/YYYY/MM/DD/rollout-*.jsonl, ~/.codex/log/, ~/.codex/config.toml |
| Gemini CLI | ~/.gemini/tmp/<project-hash>/logs.json, ~/.gemini/settings.json, checkpoint dirs under ~/.gemini/tmp/ |
| Cursor | %APPDATA%\Cursor\User\workspaceStorage\<id>\state.vscdb (SQLite — chat and composer history), %APPDATA%\Cursor\User\globalStorage\ |
| GitHub Copilot (VS Code) | %APPDATA%\Code\User\workspaceStorage\<id>\state.vscdb, chat sessions in globalStorage |
| Warp | %LOCALAPPDATA%\warp\, ~/.warp/ (SQLite and log artifacts) |
| Windsurf | ~/.codeium/windsurf/, %APPDATA%\Windsurf\User\workspaceStorage\ |
| Qwen Code | ~/.qwen/ (session and settings artifacts, mirroring Gemini CLI structure) |
| OpenCode | ~/.local/share/opencode/ (storage/message and session JSON files), project-local .opencode/ state |
| Hermes | Emerging — validate agent config directory under user profile; treat ~/.hermes and %APPDATA%\hermes as collection candidates |
Collection note: The JSONL transcripts from Claude Code and Codex are gold. Each line is a timestamped event — user prompt, assistant response, tool invocation (including full command lines), and tool results. Cursor and Copilot history requires SQLite parsing of state.vscdb (ItemTable and chat-related keys). Build these paths into your triage collection scripts now, not during the next breach.
Exploitation and Abuse Status
This is not a single-CVE story — it is an attack-surface and forensic-readiness story. The active, observed abuse patterns in 2025–2026 include:
- Prompt injection via repository content causing agents to run attacker-influenced commands (e.g., malicious
README.md, poisoned issue text, or compromised MCP server responses). - Credential and secret harvesting where agents with broad file access read
.envfiles, SSH keys, or cloud credential files and the content transits to external LLM APIs. - Supply-chain exposure via agent-generated or agent-modified code committed without meaningful human review.
- Attacker use of agents during hands-on intrusions — threat actors increasingly use local AI tooling to accelerate post-exploitation, blending malicious commands into agent session noise.
No CISA KEV entry applies here; the risk is behavioral and architectural, and the mitigation is detection, inventory, and policy.
Detection & Response
Sigma Rules
These rules target the two highest-fidelity behaviors: agent processes spawning shells/interpreters (including attacker- or injection-driven execution), and creation of agent transcript directories on systems where agents aren't sanctioned. Tune the agent binary list to your approved inventory.
---
title: AI Coding Agent Spawning Shell or Interpreter
id: 3f8a1c74-2b6e-4d91-a7c3-9e0f5b2d8a41
status: experimental
description: Detects known AI coding agent processes spawning command shells or script interpreters. Agent tool-call execution is expected in dev environments, but this lineage is the primary indicator of prompt-injection-driven execution and attacker abuse of local agents.
references:
- https://isc.sans.edu/diary/rss/33410
- https://attack.mitre.org/techniques/T1059/
author: Security Arsenal
date: 2026/01/15
tags:
- attack.execution
- attack.t1059
logsource:
category: process_creation
product: windows
detection:
selection_parent:
ParentImage|endswith:
- '\claude.exe'
- '\codex.exe'
- '\gemini.exe'
- '\opencode.exe'
- '\qwen.exe'
- '\hermes.exe'
- '\Cursor.exe'
- '\Windsurf.exe'
- '\Warp.exe'
- '\node.exe'
selection_child:
Image|endswith:
- '\cmd.exe'
- '\powershell.exe'
- '\pwsh.exe'
- '\wscript.exe'
- '\cscript.exe'
- '\mshta.exe'
- '\curl.exe'
- '\certutil.exe'
- '\bitsadmin.exe'
filter_claude_code:
ParentImage|endswith: '\node.exe'
ParentCommandLine|contains:
- 'claude'
- 'codex'
- 'gemini'
- 'opencode'
condition: selection_parent and selection_child and not filter_claude_code
falsepositives:
- Legitimate agent tool-call execution during approved development work
level: medium
---
title: AI Agent Transcript Directory Creation on Endpoint
id: 91c2e5b8-4a7f-4c36-b8d2-6f1a9e3c5d07
status: experimental
description: Detects creation of AI coding agent configuration and transcript directories. Useful for shadow-AI discovery on systems where AI agents are not sanctioned, and for establishing first-seen timelines during IR.
references:
- https://isc.sans.edu/diary/rss/33410
- https://attack.mitre.org/techniques/T1074/
author: Security Arsenal
date: 2026/01/15
tags:
- attack.collection
- attack.t1074
logsource:
category: file_event
product: windows
detection:
selection:
TargetFilename|contains:
- '\.claude\projects\'
- '\.codex\sessions\'
- '\.gemini\tmp\'
- '\.qwen\'
- '\.local\share\opencode\'
- '\.codeium\windsurf\'
- '\.hermes\'
selection_ext:
TargetFilename|endswith:
- '.jsonl'
- '.json'
condition: selection and selection_ext
falsepositives:
- Approved developer use of AI coding assistants
level: low
---
title: AI Agent Access to Credential and Secret Files
id: b47d3f29-8c1a-4e52-9d6b-2a8f4c7e1935
status: experimental
description: Detects AI agent processes accessing SSH keys, cloud credentials, or environment secret files. A primary exfiltration and over-permissioning indicator — agent tool calls reading secrets means that content likely transited to an external LLM API.
references:
- https://isc.sans.edu/diary/rss/33410
- https://attack.mitre.org/techniques/T1552/
author: Security Arsenal
date: 2026/01/15
tags:
- attack.credential_access
- attack.t1552.001
logsource:
category: file_event
product: windows
detection:
selection_paths:
TargetFilename|contains:
- '\.ssh\id_'
- '\.aws\credentials'
- '\.azure\'
- '\.config\gcloud\'
- '\.env'
- '\.npmrc'
- '\.netrc'
selection_ext:
TargetFilename|endswith:
- '\id_rsa'
- '\id_ed25519'
- 'credentials'
- '.env'
condition: selection_paths or (selection_ext and selection_paths)
falsepositives:
- Legitimate developer tooling and backup software accessing these paths
level: high
KQL — Microsoft Sentinel / Defender
This hunt identifies AI agent processes and their child processes across the fleet, with a rollup by device and agent so you can baseline approved usage and surface anomalies. The second-stage filter highlights high-risk child processes (LOLBins, downloaders, scripting hosts).
let AgentProcesses = dynamic(["claude.exe","codex.exe","gemini.exe","opencode.exe","qwen.exe","hermes.exe","Cursor.exe","Windsurf.exe","Warp.exe"]);
let HighRiskChildren = dynamic(["powershell.exe","pwsh.exe","cmd.exe","wscript.exe","cscript.exe","mshta.exe","curl.exe","certutil.exe","bitsadmin.exe","rundll32.exe","regsvr32.exe"]);
let AgentParents = DeviceProcessEvents
| where TimeGenerated > ago(14d)
| where FileName in~ (AgentProcesses)
or (FileName =~ "node.exe" and ProcessCommandLine has_any ("claude","codex","gemini-cli","opencode","qwen-code"))
| project AgentDeviceId = DeviceId, AgentDevice = DeviceName, AgentProcessId = ProcessId, AgentName = FileName, AgentCmd = ProcessCommandLine, AgentUser = AccountName, AgentTime = TimeGenerated;
DeviceProcessEvents
| where TimeGenerated > ago(14d)
| join kind=inner AgentParents on $left.DeviceId == $right.AgentDeviceId and $left.InitiatingProcessId == $right.AgentProcessId
| extend RiskFlag = iff(FileName in~ (HighRiskChildren), "HIGH", "INFO")
| summarize FirstSeen = min(TimeGenerated), LastSeen = max(TimeGenerated), ChildCount = count(), HighRiskCount = countif(RiskFlag == "HIGH"), SampleCommands = make_set(strcat(FileName, " | ", ProcessCommandLine), 5) by AgentDevice, AgentName, AgentUser, RiskFlag
| order by HighRiskCount desc, ChildCount desc
// Stage 2: Hunt transcript file creation for shadow-AI discovery and IR timeline anchoring
DeviceFileEvents
| where TimeGenerated > ago(30d)
| where FolderPath has_any ("\\.claude\\projects\\","\\.codex\\sessions\\","\\.gemini\\tmp\\","\\.qwen\\","\\opencode\\storage\\","\\.codeium\\windsurf\\","\\.hermes\\")
| summarize FirstSeen = min(TimeGenerated), LastSeen = max(TimeGenerated), FileCount = count(), SamplePaths = make_set(FolderPath, 3) by DeviceName, InitiatingProcessFileName, InitiatingProcessAccountName
| order by FirstSeen asc
Velociraptor VQL
Use this artifact during triage to enumerate installed agents and stage transcript artifacts for collection. Run it fleet-wide to build your shadow-AI inventory in one pass.
-- Hunt: enumerate AI coding agent artifacts and transcripts across user profiles
LET agent_globs = {
"ClaudeCode": "C:/Users/*/.claude/projects/**/*.jsonl",
"Codex": "C:/Users/*/.codex/sessions/**/*.jsonl",
"GeminiCLI": "C:/Users/*/.gemini/tmp/*/logs.json",
"QwenCode": "C:/Users/*/.qwen/**",
"OpenCode": "C:/Users/*/.local/share/opencode/**",
"Windsurf": "C:/Users/*/.codeium/windsurf/**",
"Hermes": "C:/Users/*/.hermes/**"
}
SELECT Agent, FullPath, Size, Mtime
FROM foreach(row=agent_globs,
query={
SELECT _value AS Agent, FullPath, Size, Mtime
FROM glob(globs=_value)
})
ORDER BY Mtime DESC
-- Hunt: running AI agent processes and their child processes
SELECT Pid, Ppid, Name, Exe, CommandLine, Username, CreateTime,
get_member(field="Name", item=process_memory(pid=Ppid)) AS ParentName
FROM pslist()
WHERE Name =~ '(?i)claude|codex|gemini|opencode|qwen|hermes|cursor|windsurf|warp'
OR CommandLine =~ '(?i)claude|codex-cli|gemini-cli|opencode|qwen-code'
Remediation / Inventory Script
Deploy this PowerShell script via your RMM, Intune, or GPO startup script to audit AI agent presence per endpoint. It enumerates known install locations and transcript directories, flags unexpected finds, and writes a JSON report for central collection.
# AI Agent Inventory & Audit Script — run as SYSTEM or admin for full user-profile coverage
# Output: JSON report per endpoint at C:\ProgramData\SecurityOps\AIAgentAudit-<hostname>.json
$reportDir = 'C:\ProgramData\SecurityOps'
New-Item -ItemType Directory -Path $reportDir -Force | Out-Null
# Define approved agents for YOUR environment — anything outside this list is flagged
$approvedAgents = @('claude', 'copilot')
$agentArtifacts = @{
'ClaudeCode' = @('.claude\projects', '.claude.json')
'Codex' = @('.codex\sessions', '.codex\config.toml')
'GeminiCLI' = @('.gemini\tmp', '.gemini\settings.json')
'Cursor' = @('AppData\Roaming\Cursor\User\workspaceStorage')
'Copilot' = @('AppData\Roaming\Code\User\workspaceStorage')
'Warp' = @('AppData\Local\warp')
'Windsurf' = @('.codeium\windsurf')
'QwenCode' = @('.qwen')
'OpenCode' = @('.local\share\opencode')
'Hermes' = @('.hermes', 'AppData\Roaming\hermes')
}
$findings = @()
Get-ChildItem 'C:\Users' -Directory -ErrorAction SilentlyContinue | ForEach-Object {
$userProfile = $_.FullName
foreach ($agent in $agentArtifacts.GetEnumerator()) {
foreach ($relPath in $agent.Value) {
$full = Join-Path $userProfile $relPath
if (Test-Path $full) {
$item = Get-Item $full -ErrorAction SilentlyContinue
$agentKey = $agent.Key.ToLower()
$findings += [PSCustomObject]@{
Agent = $agent.Key
UserProfile = $userProfile
Path = $full
LastWrite = $item.LastWriteTime
Approved = ($approvedAgents | Where-Object { $agentKey -like "*$_*" }).Count -gt 0
}
}
}
}
}
# Check running agent processes right now
$running = Get-Process -ErrorAction SilentlyContinue |
Where-Object { $_.Name -match 'claude|codex|gemini|opencode|qwen|hermes|cursor|windsurf|warp' } |
Select-Object Name, Id, Path
$report = [PSCustomObject]@{
Hostname = $env:COMPUTERNAME
ScanTime = (Get-Date).ToString('o')
ArtifactsFound = $findings
UnauthorizedCount = ($findings | Where-Object { -not $_.Approved }).Count
RunningAgents = $running
}
$outFile = Join-Path $reportDir "AIAgentAudit-$env:COMPUTERNAME.json"
$report | ConvertTo-Json -Depth 4 | Out-File $outFile -Encoding utf8
Write-Output "Audit complete: $($findings.Count) artifacts, $($report.UnauthorizedCount) unauthorized -> $outFile"
For Linux/macOS developer fleets, this Bash equivalent covers the primary paths:
#!/usr/bin/env bash
# AI agent artifact audit for Linux/macOS — run via your MDM or config management
REPORT="/var/log/ai-agent-audit-$(hostname)-$(date +%F).json"
echo "{\"hostname\":\"$(hostname)\",\"findings\":[" > "$REPORT"
first=1
for home in /home/* /Users/* /root; do
for path in ".claude/projects" ".codex/sessions" ".gemini/tmp" ".qwen" \
".local/share/opencode" ".codeium/windsurf" ".hermes" ".warp"; do
full="$home/$path"
if [ -e "$full" ]; then
[ $first -eq 0 ] && echo "," >> "$REPORT"
first=0
mtime=$(stat -c '%y' "$full" 2>/dev/null || stat -f '%Sm' "$full" 2>/dev/null)
echo "{\"user\":\"$home\",\"path\":\"$full\",\"mtime\":\"$mtime\"}" >> "$REPORT"
fi
done
done
echo "]}" >> "$REPORT"
# Flag running agents
ps -eo user,comm,args | grep -iE 'claude|codex|gemini|opencode|qwen|hermes' | grep -v grep
Remediation
There is no patch for this problem — it requires governance, architecture, and forensic readiness. Priority actions:
- Establish an approved-agent inventory. Use the audit scripts above fleet-wide this week. Anything outside your approved list is shadow AI: either sanction it through a review process or remove it via your software management tooling.
- Constrain agent permissions. Run agents with least privilege: dedicated service accounts or sandboxed dev containers, no standing access to production credentials, and
.env/SSH/cloud credential stores outside agent-readable scope. Where supported, configure agent allowlists for tool calls (e.g., Claude Code permission modes, Codex sandboxing flags like--sandbox read-onlyor workspace-write scoping). - Update your IR playbooks and triage collections. Add the artifact table above to your forensic acquisition checklists. JSONL transcripts and
state.vscdbfiles must be captured before reimaging developer workstations. SANS FOR577's updated day-5 material is a solid reference if your team needs structured training. - Deploy the detections above. The parent-child process lineage rule is your tripwire for prompt-injection-driven execution; the transcript-creation rule builds your shadow-AI detection over time.
- Govern MCP servers. Treat MCP integrations as third-party software: inventory them, pin versions, and block unapproved endpoints. A compromised MCP server is a direct injection path into every connected agent.
- Log and review agent network egress. Agents beacon to vendor APIs (anthropic.com, openai.com, generativelanguage.googleapis.com, etc.). Unexpected destinations from agent processes warrant investigation.
- Set policy for agent-generated code. Require human review and signed commits for agent-authored changes to production codebases; treat agent commits as a distinct, auditable identity where your VCS supports it.
Related Resources
Security Arsenal Penetration Testing Services AlertMonitor Platform Book a SOC Assessment vulnerability-management Intel Hub
Is your security operations ready?
Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.