Security architectures are built on a set of assumptions that autonomous AI agents quietly invalidate: that identities are human, that API integrations are static and reviewed, and that an action on the network traces back to a person who can be interviewed, revoked, and held accountable. As organizations move from experimenting with a single copilot to deploying agents that delegate tasks to other agents, invoke tools, and chain actions across SaaS, cloud, and internal systems, those assumptions collapse.
The scenario is already operational in most enterprises: a primary orchestration agent receives a business goal, decomposes it, and delegates subtasks to secondary agents. Those secondary agents query production databases, call internal APIs, open tickets, send emails, and in some cases provision or modify cloud resources — all without a human in the loop. Every hop in that chain is an authentication event, an authorization decision, and a potential abuse path. When I run tabletop exercises with clients today, the question that consistently goes unanswered is the most basic one: whose identity did that action happen under?
If the answer is "a shared service account" or "the API key we embedded in the agent framework," you have an accountability and detection gap that adversaries will find before your SOC does. This post breaks down the agent-to-agent attack surface, gives you concrete detection content for the behaviors that matter, and lays out a remediation path that doesn't require blocking adoption.
Technical Analysis: Why Agent-to-Agent Is a New Attack Surface
What Breaks in Traditional Architectures
Classic enterprise security assumes a one-to-one relationship between an authenticated session and a responsible party. Agentic systems break this in four specific ways:
1. Delegation chains flatten privilege. When Agent A delegates to Agent B, Agent B typically inherits Agent A's token, service account, or broad API scope. There is no attenuation — the delegated agent carries the full authority of the delegator, and often the full authority of the human who kicked off the workflow. An instruction injected into Agent B's context (via a poisoned document, email, or web page it was asked to summarize) now executes with Agent A's — and the originating user's — privileges.
2. Non-human identities (NHIs) proliferate unmanaged. Every agent, tool integration, MCP (Model Context Protocol) server connection, and webhook is a non-human identity. In most environments we assess, NHIs already outnumber human users by an order of magnitude, and agent deployments are accelerating that ratio. These identities are frequently created by developers, granted broad OAuth scopes (e.g., Mail.ReadWrite, Sites.FullControl.All, or * IAM bindings) "to make it work," and never reviewed again.
3. Behavioral baselines don't exist. Your UEBA and SIEM detections were tuned on human patterns: business-hours activity, bounded request rates, predictable application sequences. Agents operate at machine speed, around the clock, with bursty and highly variable behavior. Either your detections drown in noise and get disabled, or agents get blanket exclusions — which adversaries then hide behind.
4. The confused deputy problem returns at scale. An agent with legitimate credentials can be manipulated — through indirect prompt injection — into performing unauthorized actions against systems it is authorized to reach. The authentication logs will show a perfectly valid token used in a perfectly normal way. The attack lives entirely in the semantic layer: what the agent was told to do, not how it authenticated.
Attacker Opportunity Set
From the red team side of my work, the realistic abuse paths against agent-to-agent architectures are:
- Indirect prompt injection → tool abuse. Content ingested by an agent (emails, tickets, documents, web pages) carries instructions that redirect the agent's tool calls — exfiltrating data to an attacker-controlled endpoint, modifying records, or invoking downstream agents with crafted tasks.
- Token and credential theft from agent runtimes. Agents hold long-lived API keys, OAuth refresh tokens, or service account credentials, often stored in environment variables, config files, or secrets stores with weak scoping. Compromising the agent host or its orchestration layer yields high-value, poorly monitored credentials.
- Delegation spoofing. Where agent-to-agent trust is implicit (shared queues, unauthenticated internal endpoints, unsigned task messages), an attacker who can write to the message path can impersonate a legitimate delegating agent.
- Scope creep via tool registration. Dynamically registered tools and MCP servers expand what an agent can do without a corresponding access review. A malicious or compromised tool server becomes a privileged execution path.
Exploitation Status
This is an emerging attack surface, not a single CVE — and no CVE is associated with this advisory. Exploitation is real but early-stage: security research through 2025 has demonstrated indirect prompt injection leading to data exfiltration and tool abuse against production agent frameworks, and we're seeing threat actors begin probing AI integrations during intrusions. The defensive window is open now, before agent-specific tradecraft matures. Treat this with the same urgency you gave identity hardening after the first wave of token-theft attacks against cloud tenants.
Detection & Response
Detection for agentic environments centers on three observable behaviors: non-human identities acting outside their behavioral baseline, agent runtimes spawning unexpected child processes or network connections, and privilege grants to application identities that expand the agent attack surface. The rules below are tuned to fire on deviations worth investigating — not on every agent action.
Sigma Rules
---
title: Agent Runtime Spawning Interactive Shell or Command Interpreter
id: 3f8a2c41-9b1e-4d72-a6c5-8e2f1b3d4a97
status: experimental
description: Detects common AI agent runtime processes (Python, Node.js) spawning interactive shells or command interpreters. Agent frameworks executing tasks should invoke defined tools and APIs, not drop to cmd, PowerShell, or bash — a strong indicator of prompt-injection-driven tool abuse or runtime compromise.
references:
- https://www.rapid7.com/blog/post/ai-securing-agent-to-agent-communication-next-identity-frontier
- https://attack.mitre.org/techniques/T1059/
author: Security Arsenal
date: 2026/02/10
tags:
- attack.execution
- attack.t1059
logsource:
category: process_creation
product: windows
detection:
selection_parent:
ParentImage|endswith:
- '\python.exe'
- '\python3.exe'
- '\node.exe'
- '\deno.exe'
selection_child:
Image|endswith:
- '\cmd.exe'
- '\powershell.exe'
- '\pwsh.exe'
- '\wscript.exe'
- '\cscript.exe'
- '\mshta.exe'
- '\rundll32.exe'
- '\certutil.exe'
- '\bitsadmin.exe'
- '\curl.exe'
- '\wget.exe'
filter_known_paths:
ParentImage|contains:
- '\nodejs\node_modules\.bin\'
condition: selection_parent and selection_child and not filter_known_paths
falsepositives:
- Agent frameworks with a legitimate shell-execution tool enabled (review and explicitly allowlist the exact command patterns, do not blanket-exclude)
- Developer workstations running local agent tooling
level: high
---
title: Highly Privileged OAuth Scope Granted to Application or Service Principal
id: 7c1d4e82-2a5f-4b93-9d18-5f6a7c8e9b01
status: experimental
description: Detects consent grants of high-impact Microsoft Graph / Azure AD scopes to service principals and app registrations. Agentic integrations are frequently over-provisioned with broad scopes (Mail.ReadWrite, Sites.FullControl.All, Directory access) that turn a compromised agent into a tenant-wide foothold.
references:
- https://www.rapid7.com/blog/post/ai-securing-agent-to-agent-communication-next-identity-frontier
- https://attack.mitre.org/techniques/T1550/001/
author: Security Arsenal
date: 2026/02/10
tags:
- attack.persistence
- attack.privilege_escalation
- attack.t1550.001
logsource:
product: azure
service: auditlogs
detection:
selection_operation:
OperationName:
- 'Consent to application'
- 'Add delegated permission grant'
- 'Add app role assignment to service principal'
selection_scope:
TargetResources|contains:
- 'Mail.ReadWrite'
- 'Mail.Send'
- 'Sites.FullControl.All'
- 'Files.ReadWrite.All'
- 'Directory.ReadWrite.All'
- 'RoleManagement.ReadWrite.Directory'
- 'Application.ReadWrite.All'
- 'full_access_as_app'
condition: selection_operation and selection_scope
falsepositives:
- Approved onboarding of enterprise applications — route through change control and match against approved app IDs
level: high
---
title: Non-Human Identity Activity Outside Baseline Hours or Source
id: 9e4b6f13-7d2a-4c85-b3e6-1a9d2f5c7e48
status: experimental
description: Detects service account authentication to interactive or sensitive resources from workstations or unusual sources. Agent and service identities should authenticate from fixed infrastructure; logons from user workstations or new hosts indicate credential theft from an agent runtime or secrets store.
references:
- https://www.rapid7.com/blog/post/ai-securing-agent-to-agent-communication-next-identity-frontier
- https://attack.mitre.org/techniques/T1078/004/
author: Security Arsenal
date: 2026/02/10
tags:
- attack.initial_access
- attack.defense_evasion
- attack.t1078.004
logsource:
category: authentication
product: windows
detection:
selection:
User|contains:
- 'svc-'
- 'svc_'
- '-svc'
- 'agent-'
- 'automation'
LogonType:
- 2
- 10
condition: selection
falsepositives:
- Administrators interactively troubleshooting under a service account (bad practice — detect and correct it)
- Misnamed human accounts matching the service account naming convention
level: medium
A note on tuning: rule one is your highest-signal rule. Legitimate agent frameworks that expose a shell tool should have that tool disabled in production — if the business insists on it, allowlist exact command-line patterns, never the parent process. Rule three assumes a disciplined naming convention for non-human identities; if you don't have one, building it is part of the remediation below.
KQL (Microsoft Sentinel / Defender)
This hunt identifies service principals — the identity type most agent integrations use in Entra ID — authenticating to resources they've never touched before or from new IP addresses. New-resource access by an NHI is the cloud-side signature of a hijacked agent token being repurposed.
// Hunt: Service principals accessing resources or IPs outside their 14-day baseline
// Tune: exclude known automation principals after validation; alert on the rest
let lookback = 14d;
let baseline =
AADServicePrincipalSignInLogs
| where TimeGenerated between (ago(lookback*2) .. ago(lookback))
| summarize BaselineResources = make_set(ResourceDisplayName),
BaselineIPs = make_set(IPAddress)
by ServicePrincipalId, AppDisplayName;
AADServicePrincipalSignInLogs
| where TimeGenerated > ago(24h)
| where ResultType == 0
| join kind=inner baseline on ServicePrincipalId
| where not(set_has_element(BaselineResources, ResourceDisplayName))
or not(set_has_element(BaselineIPs, IPAddress))
| summarize FirstSeen = min(TimeGenerated),
NewResources = make_set_if(ResourceDisplayName, not(set_has_element(BaselineResources, ResourceDisplayName))),
NewIPs = make_set_if(IPAddress, not(set_has_element(BaselineIPs, IPAddress))),
Requests = count()
by AppDisplayName, ServicePrincipalId, AppId
| order by Requests desc;
// Hunt: Agent runtime processes making unexpected outbound network connections
// Looks for python/node processes connecting to rare external destinations
DeviceNetworkEvents
| where TimeGenerated > ago(7d)
| where InitiatingProcessFileName in~ ("python.exe", "python3.exe", "node.exe", "deno.exe", "python", "node")
| where RemoteIPType == "Public"
| summarize Connections = count(),
Ports = make_set(RemotePort),
Processes = make_set(InitiatingProcessCommandLine)
by DeviceName, RemoteUrl, RemoteIP, InitiatingProcessFileName
| where Connections < 50 // Rare destinations — high-volume endpoints are usually the LLM API itself
| order by Connections asc;
Velociraptor VQL
Use this artifact to hunt across your fleet for agent runtime processes holding network connections to unusual destinations, alongside their loaded environment — a common place to find exposed API keys during triage.
-- Hunt: Agent runtime processes with active external connections
-- Identifies python/node processes talking to non-standard destinations,
-- useful for finding compromised agent hosts or data exfiltration channels
LET proc_names = {'python', 'python3', 'node', 'deno', 'uvicorn', 'gunicorn'}
SELECT Pid,
Name,
CommandLine,
Exe,
Username,
CreateTime,
netstat().RemoteIP AS RemoteIP,
netstat().RemotePort AS RemotePort,
netstat().Status AS ConnStatus
FROM pslist()
WHERE Name =~ '(?i)^(python3?|node|deno|uvicorn|gunicorn)(\\.exe)?$'
AND netstat().RemoteIP !~ '^(127\\.|10\\.|172\\.(1[6-9]|2[0-9]|3[01])\\.|192\\.168\\.|::1)'
AND netstat().Status =~ 'ESTABLISHED'
-- Hunt: Exposed API keys and tokens in agent process environments
-- Agent frameworks commonly receive credentials via environment variables;
-- triage artifact to identify which agent identities are exposed per host
SELECT Pid,
Name,
CommandLine,
Username,
environ(key=EnvVar) AS EnvValue
FROM pslist()
LET EnvVar = 'OPENAI_API_KEY'
WHERE Name =~ '(?i)(python3?|node|deno)'
AND environ(key=EnvVar)
Note: extend the second artifact's variable list (ANTHROPIC_API_KEY, AZURE_CLIENT_SECRET, AWS_SECRET_ACCESS_KEY) per your stack — and treat any hit as a finding in itself. Long-lived secrets in process environments are exactly what this architecture needs to eliminate.
Remediation: Making Agents First-Class Identities
There is no patch for an architectural gap. The remediation is a program, and it needs to start before your agent deployments outpace your ability to inventory them.
1. Inventory every non-human identity — now. Enumerate service principals, app registrations, service accounts, API keys, and agent framework identities across cloud and on-prem. For Entra ID, the PowerShell script below produces the baseline audit. You cannot baseline behavior for identities you don't know exist.
# Audit non-human identities: app registrations with credentials and high-privilege scopes
# Requires: Microsoft.Graph PowerShell SDK, Application.Read.All and AuditLog.Read.All
Connect-MgGraph -Scopes "Application.Read.All","Directory.Read.All","AuditLog.Read.All"
# Find service principals with credentials (secrets or certificates) and their expiry
$report = foreach ($sp in (Get-MgServicePrincipal -All)) {
foreach ($cred in $sp.PasswordCredentials) {
[PSCustomObject]@{
DisplayName = $sp.DisplayName
AppId = $sp.AppId
CredType = 'Secret'
ExpiresOn = $cred.EndDateTime
DaysToExpiry = [math]::Round((New-TimeSpan -Start (Get-Date) -End $cred.EndDateTime).TotalDays)
SignInEnabled = $sp.AccountEnabled
}
}
foreach ($cert in $sp.KeyCredentials) {
[PSCustomObject]@{
DisplayName = $sp.DisplayName
AppId = $sp.AppId
CredType = 'Certificate'
ExpiresOn = $cert.EndDateTime
DaysToExpiry = [math]::Round((New-TimeSpan -Start (Get-Date) -End $cert.EndDateTime).TotalDays)
SignInEnabled = $sp.AccountEnabled
}
}
}
# Flag long-lived credentials (>180 days) — prime candidates for rotation to short-lived auth
$report | Where-Object { $_.DaysToExpiry -gt 180 } | Sort-Object DaysToExpiry -Descending |
Export-Csv -Path ".\NHI_LongLived_Credentials.csv" -NoTypeInformation
# Enumerate app role assignments (application permissions) per service principal
$grants = foreach ($sp in (Get-MgServicePrincipal -All)) {
Get-MgServicePrincipalAppRoleAssignment -ServicePrincipalId $sp.Id -All |
Select-Object @{n='ServicePrincipal';e={$sp.DisplayName}},
AppRoleId, ResourceDisplayName, PrincipalType
}
$grants | Export-Csv -Path ".\NHI_AppRole_Grants.csv" -NoTypeInformation
Write-Output "Exported NHI credential inventory and app role grants. Review for over-privileged and long-lived entries."
2. Issue agents their own identities — never shared, never inherited. Each agent gets a distinct identity with its own credentials, and delegation between agents must use token exchange or on-behalf-of flows that attenuate scope at each hop rather than passing the parent's token. If Agent B only needs to read one SharePoint site, its token should say exactly that — even if Agent A can read all of them.
3. Eliminate long-lived secrets from agent runtimes. Move to workload identity federation, managed identities, or SPIFFE/SPIRE-issued SVIDs with lifetimes measured in minutes. Mutual TLS between agents provides both authentication and integrity for agent-to-agent messages — it defeats delegation spoofing on the message path.
4. Enforce least privilege on tool access, not just API scopes. Inventory every tool and MCP server an agent can invoke. Disable shell-execution and arbitrary code tools in production agents. Require explicit allowlists for outbound destinations from agent runtimes — an agent summarizing documents has no business initiating arbitrary HTTPS to the internet.
5. Build behavioral baselines per agent identity. Log every tool invocation, every inter-agent delegation, and every resource access keyed to the agent's own identity. Establish per-agent baselines (resources, request rates, hours, destinations) in your SIEM — the KQL above is a starting pattern. An agent deviating from its own baseline is your earliest indicator of prompt injection or token theft.
6. Keep humans in the loop for irreversible actions. Wire transfers, permission changes, mass communications, production deployments, and deletions should require human approval regardless of how many agents are in the chain. This is a policy control enforced at the tool/API layer — not a prompt instruction, which an injection attack can simply override.
7. Red-team the agent layer. Add indirect prompt injection and tool-abuse scenarios to your penetration test scope. Test whether a poisoned document can redirect your agents, whether a compromised secondary agent can escalate through the delegation chain, and whether your SOC actually detects it. If your IR plan doesn't include revoking agent identities and rotating their credentials at machine speed, update it.
The Bottom Line
Agent-to-agent communication is not a future problem — it's running in your environment today, likely under identities your IAM program doesn't fully track. The organizations that get ahead of this will do three things: give every agent its own tightly scoped identity, replace inherited and long-lived credentials with short-lived attested ones, and build detection against per-agent behavioral baselines rather than human norms. The attack surface is real, the tradecraft against it is immature, and that combination is exactly when defenders should move.
Related Resources
Security Arsenal Managed SOC Services AlertMonitor Platform Book a SOC Assessment soc-mdr Intel Hub
Is your security operations ready?
Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.