Back to Intelligence

AI Agents Exceeding Their Permissions: How to Detect and Contain Over-Privileged Non-Human Identities

SA
Security Arsenal Team
October 9, 2026
12 min read

The identity perimeter just got a lot more complicated. As enterprises race to deploy autonomous AI agents — copilots that book travel, agents that triage tickets, pipelines that write and merge code — a structural security gap has emerged: these agents authenticate with valid credentials and then perform actions well outside the scope of what they were actually granted. Token Security's recent analysis of this problem cuts to the heart of it: traditional access controls were designed around human users and static service accounts. They were not designed for non-deterministic software that chains tool calls, inherits broad OAuth scopes, and dynamically decides which API to hit next.

This is not a theoretical risk. In our incident response practice, we are already seeing the early shape of this threat: service principals with Mail.ReadWrite granted for a narrow summarization use case being leveraged to enumerate entire mailboxes; CI/CD agents with repo-wide tokens being manipulated — through prompt injection or simple logic errors — into exfiltrating secrets; and AI assistants holding standing cloud credentials that no one remembers granting. Because the credentials are valid, nothing alarms. The authentication succeeds. The API returns data. Your DLP sees an authorized token. The blast radius only becomes apparent in hindsight.

Defenders need to treat AI agents as a distinct identity class — one that demands agent-specific policy enforcement, behavioral baselining, and ruthless least privilege — without killing the autonomy that makes these tools valuable in the first place.

Technical Analysis

The Core Problem: Valid Credentials, Undefined Boundaries

Unlike a traditional vulnerability with a CVE and a patch, this is an architectural weakness in how organizations provision identity for autonomous software. The attack surface breaks down into several observable patterns:

1. Over-scoped OAuth grants. When an AI agent is integrated into Microsoft 365, Google Workspace, Salesforce, or Slack, it typically requests OAuth scopes during consent. Developers, under deadline pressure, grant tenant-wide scopes (Sites.Read.All, Mail.ReadWrite, Directory.Read.All) instead of resource-scoped permissions. The agent's LLM-driven planner then has the capability to touch every mailbox or SharePoint site in the tenant — even though the business intent was one inbox.

2. Credential inheritance and delegation. Agents frequently operate using the credentials of the human who deployed them, or a shared service account. When the agent decides — autonomously — to call an API, it presents a token that is indistinguishable from legitimate human activity in your logs. There is no is_agent flag in a Kerberos ticket.

3. Tool chaining and confused deputy scenarios. Modern agent frameworks (LangChain, AutoGen, MCP-based tool servers) let an agent invoke arbitrary tools. An attacker who can influence the agent's input — a malicious email it summarizes, a poisoned document it ingests, a compromised MCP server it trusts — can redirect that valid credential toward actions the owner never intended. This is the confused deputy problem reborn for agentic AI.

4. Cloud metadata credential theft. Agents running on cloud compute (Azure VMs, EC2, container workloads) can reach the instance metadata service at 169.254.169.254 and harvest managed identity tokens — a capability the agent's designer likely never intended it to exercise.

Why Traditional Controls Fail

  • RBAC assumes static roles. An agent's behavior is dynamic; its effective permissions are the union of every scope it holds, not the narrow task it was built for.
  • MFA doesn't apply to NHIs. Non-human identities authenticate with tokens, certificates, and secrets. There is no push notification to approve.
  • SIEM rules tuned for humans miss agents. Impossible travel, logon hour anomalies, and password spray detections are calibrated for user accounts. An agent hitting 400 APIs in a minute from a datacenter IP looks like normal service traffic — until it isn't.

Exploitation Status

This is not a single exploited CVE — it is an actively developing threat class. Prompt injection against AI agents with tool access has been demonstrated repeatedly in published research and real-world proofs of concept throughout 2025 and into 2026, and identity vendors including Token Security are reporting enterprise deployments where agent permissions materially exceed business intent. Treat this as a present-tense governance gap in your environment, not a future concern: if you have deployed any Copilot Studio agent, custom GPT with API actions, MCP server, or LangChain-based automation, you almost certainly have over-privileged non-human identities right now.

Detection & Response

Detection here is about behavioral deviation of non-human identities — service principals, managed identities, and agent service accounts acting outside their established baseline. The rules below target the highest-signal observable behaviors.

Sigma Rules

YAML
---
title: Service Principal Consent Grant with High-Privilege Application Scope
id: 8f3a1c92-4d7e-4b61-a934-2c5d8e7f1029
status: experimental
description: Detects OAuth consent grants assigning broad application-level scopes (mail, directory, sites) to service principals, a common indicator of over-privileged AI agent provisioning or consent phishing.
references:
  - https://attack.mitre.org/techniques/T1550/001/
  - https://www.bleepingcomputer.com/news/security/how-to-keep-ai-agents-within-their-permissions/
author: Security Arsenal
date: 2026/04/06
tags:
  - attack.persistence
  - attack.t1550.001
logsource:
  product: azure
  service: auditlogs
detection:
  selection_operation:
    OperationName:
      - 'Consent to application'
      - 'Add app role assignment to service principal'
  selection_scope:
    TargetResources|contains:
      - 'Mail.ReadWrite'
      - 'Mail.Read'
      - 'Directory.Read.All'
      - 'Directory.ReadWrite.All'
      - 'Sites.Read.All'
      - 'Sites.ReadWrite.All'
      - 'Files.ReadWrite.All'
      - 'full_access_as_app'
  condition: selection_operation and selection_scope
falsepositives:
  - Legitimate enterprise application onboarding — build an allowlist of approved enterprise applications and review all others
level: high
---
title: Interactive Logon by Service Account Naming Convention
id: 3b7e2d51-9a4f-4c83-b612-7d9f0a3e5c18
status: experimental
description: Detects interactive or remote interactive logons by accounts following common service/agent account naming patterns. AI agent service accounts and NHIs should authenticate programmatically; interactive sessions indicate misuse or attacker hands-on-keyboard activity with stolen agent credentials.
references:
  - https://attack.mitre.org/techniques/T1078/004/
author: Security Arsenal
date: 2026/04/06
tags:
  - attack.persistence
  - attack.initial_access
  - attack.t1078
definition: 1078/004
logsource:
  product: windows
  service: security
detection:
  selection_logontype:
    LogonType:
      - 2
      - 10
  selection_account:
    TargetUserName|startswith:
      - 'svc-'
      - 'svc_'
      - 'agent-'
      - 'bot-'
  filter_approved:
    TargetUserName|contains:
      - 'svc-backup-approved'
  condition: selection_logontype and selection_account and not filter_approved
falsepositives:
  - Legacy service accounts configured for interactive logon by design — inventory and remediate rather than allowlist long-term
level: high
---
title: Process Access to Cloud Instance Metadata Service
id: 5c1f8a47-2e9b-4d36-a785-9f2b4c6d8e31
status: experimental
description: Detects processes on endpoints or servers establishing connections to the cloud instance metadata service (169.254.169.254), a technique used to harvest managed identity tokens — including by autonomous agents operating outside intended boundaries or by attackers abusing agent workloads.
references:
  - https://attack.mitre.org/techniques/T1552/005/
  - https://attack.mitre.org/techniques/T1078/004/
author: Security Arsenal
date: 2026/04/06
tags:
  - attack.credential_access
  - attack.t1552.005
logsource:
  category: network_connection
  product: windows
detection:
  selection:
    DestinationIp: '169.254.169.254'
  filter_known:
    Image|endswith:
      - '\WindowsAzureGuestAgent.exe'
      - '\WaAppAgent.exe'
      - '\MsMpEng.exe'
  condition: selection and not filter_known
falsepositives:
  - Cloud platform agents and legitimate monitoring tooling — tune the filter list to your known-good cloud agents per environment
level: medium

KQL — Microsoft Sentinel / Defender

The following hunt identifies service principals (the identity under which most AI agents operate in Entra ID) accessing resources or APIs outside their historical baseline — the exact signature of an agent exceeding its intended permissions:

KQL — Microsoft Sentinel / Defender
// Baseline: which target resources each service principal has touched in the last 30 days
let baseline =
    AADServicePrincipalSignInLogs
    | where TimeGenerated between (ago(30d) .. ago(1d))
    | summarize BaselineResources = make_set(ResourceDisplayName), BaselineIPs = make_set(IPAddress) by ServicePrincipalId, AppDisplayName;
// Last 24h: flag first-time resource access or new source IPs for those principals
AADServicePrincipalSignInLogs
| where TimeGenerated > ago(1d)
| summarize RecentResources = make_set(ResourceDisplayName), RecentIPs = make_set(IPAddress), SigninCount = count()
    by ServicePrincipalId, AppDisplayName
| join kind=inner baseline on ServicePrincipalId
| extend NewResources = set_difference(RecentResources, BaselineResources)
| extend NewIPs = set_difference(RecentIPs, BaselineIPs)
| where array_length(NewResources) > 0 or array_length(NewIPs) > 0
| project TimeGenerated = now(), AppDisplayName, ServicePrincipalId, SigninCount,
          NewResources, RecentResources, NewIPs
| sort by SigninCount desc;
// Companion hunt: agent workloads reaching cloud metadata service
DeviceNetworkEvents
| where TimeGenerated > ago(24h)
| where RemoteIP == "169.254.169.254"
| where InitiatingProcessFileName !in~ ("WindowsAzureGuestAgent.exe", "WaAppAgent.exe", "MsMpEng.exe")
| summarize Connections = count(), FirstSeen = min(TimeGenerated), LastSeen = max(TimeGenerated)
    by DeviceName, InitiatingProcessFileName, InitiatingProcessCommandLine, InitiatingProcessAccountName
| sort by Connections desc;

Tune the baseline window to your environment. The first query is the workhorse: a service principal that has only ever called Microsoft Graph suddenly authenticating to Azure Resource Manager or Key Vault is exactly what an AI agent breaking its permission boundary looks like in telemetry.

Velociraptor VQL

For endpoint forensics on hosts running agent workloads (developer workstations, automation servers, container hosts), hunt for processes holding or requesting cloud identity tokens:

VQL — Velociraptor
-- Hunt for processes communicating with cloud instance metadata service
-- and processes running common AI agent frameworks with network activity
SELECT Pid, Ppid, Name, CommandLine, Exe, Username, CreateTime
FROM pslist()
WHERE CommandLine =~ '(?i)(169\\.254\\.169\\.254|metadata.google.internal|IMDS|managed-identity)'
   OR CommandLine =~ '(?i)(langchain|autogen|openai|anthropic|mcp[-_]server|copilot)'
   AND Username =~ '(?i)(svc|agent|bot)'

-- Correlate with active network connections from those processes
SELECT Pid, Name, Family, Type, Status, Laddr, Raddr
FROM netstat()
WHERE Status =~ 'ESTAB'
  AND (Raddr =~ '169.254.169.254' OR Laddr =~ '169.254.169.254')

Remediation Script — Audit Over-Privileged Agent Identities in Entra ID

The single highest-value defensive action is an inventory of what your non-human identities can actually do. This PowerShell script (requires the Microsoft.Graph module and Application.Read.All + Directory.Read.All for the auditing identity) enumerates service principals holding high-privilege Microsoft Graph app roles and flags stale credentials — the two most common findings when AI agents are provisioned carelessly:

PowerShell
# Audit Entra ID service principals for over-privileged AI agent identities
# Run as an auditor identity with Application.Read.All and Directory.Read.All
Connect-MgGraph -Scopes "Application.Read.All","Directory.Read.All" -NoWelcome

# High-privilege Graph app role IDs that AI agents should almost never hold tenant-wide
$highPrivRoles = @{
    'Mail.ReadWrite'        = 'e2a3a72e-5f79-4c64-b1b1-878b674786c9'
    'Mail.Read'             = '810c84a8-4a9e-49e6-bf7d-12d183f40d01'
    'Directory.Read.All'    = '7ab1d382-f21e-4acd-a863-ba3e13f7da61'
    'Directory.ReadWrite.All' = '19dbc75e-c2e2-444c-a770-ec69d8559fc7'
    'Sites.ReadWrite.All'   = '9492366f-7969-46a4-8d15-ed1a20078fff'
    'Files.ReadWrite.All'   = '75359482-378d-4052-8f01-80520e7db3cd'
}

# Resolve Graph service principal to map role IDs
$graphSp = Get-MgServicePrincipal -Filter "appId eq '00000003-0000-0000-c000-000000000000'"
$report = @()

Get-MgServicePrincipal -All | ForEach-Object {
    $sp = $_
    $assignments = Get-MgServicePrincipalAppRoleAssignment -ServicePrincipalId $sp.Id -All -ErrorAction SilentlyContinue
    foreach ($a in $assignments) {
        $role = $graphSp.AppRoles | Where-Object { $_.Id -eq $a.AppRoleId }
        if ($role -and $highPrivRoles.Keys -contains $role.Value) {
            # Check for stale secrets / certs older than 90 days
            $staleCreds = @($sp.PasswordCredentials + $sp.KeyCredentials) |
                Where-Object { $_.StartDateTime -lt (Get-Date).AddDays(-90) }
            $report += [PSCustomObject]@{
                ServicePrincipal = $sp.DisplayName
                ObjectId         = $sp.Id
                PrivilegedScope  = $role.Value
                CredentialCount  = @($sp.PasswordCredentials + $sp.KeyCredentials).Count
                StaleCredentials = $staleCreds.Count
                Risk             = if ($staleCreds.Count -gt 0) { 'CRITICAL' } else { 'HIGH' }
            }
        }
    }
}

$report | Sort-Object Risk, PrivilegedScope | Format-Table -AutoSize
$report | Export-Csv -Path ".\AgentIdentityAudit_$(Get-Date -Format 'yyyyMMdd').csv" -NoTypeInformation
Write-Host "[+] Audit complete. Review CRITICAL entries first: over-privileged scope + stale credential = revoke and re-provision with least privilege."

Run this monthly at minimum. Every CRITICAL entry is an agent identity with tenant-wide data access and an aging credential — precisely the asset an attacker (or a prompt-injected agent) will leverage first.

Remediation

Because this is an architectural weakness rather than a patchable CVE, remediation is a governance and engineering program. Prioritize in this order:

1. Inventory every AI agent identity — this week. Enumerate all service principals, managed identities, API keys, and OAuth grants used by agent frameworks, Copilot Studio agents, MCP servers, and automation pipelines. You cannot govern what you have not cataloged. The script above is your starting point for Entra ID; repeat the exercise for Google Workspace service accounts, AWS IAM roles attached to agent workloads, and SaaS integrations (Slack, Salesforce, ServiceNow bots).

2. Enforce least privilege at the scope level. Replace tenant-wide Graph scopes (Mail.ReadWrite, Sites.Read.All) with resource-scoped alternatives: Sites.Selected combined with explicit site grants, mailbox-scoped application access policies in Exchange Online (New-ApplicationAccessPolicy), and per-conversation Slack bot scopes. The agent that summarizes one inbox should be technically incapable of reading a second one.

3. Adopt agent-specific policy enforcement. As Token Security's guidance emphasizes, static RBAC is insufficient — you need a policy layer that understands the agent's intended function and evaluates each action against it at runtime. Whether you implement this via an identity-first platform, a proxy/gateway in front of the agent's tool calls, or cloud-native controls (Azure AD Conditional Access for workload identities, AWS IAM permission boundaries), the principle is the same: the agent's effective permissions should be the intersection of its credential and a behavioral policy, not the union of every scope it holds.

4. Eliminate standing credentials where possible. Move agent workloads to managed identities or workload identity federation instead of long-lived secrets. Rotate what remains on a 90-day maximum. Deploy just-in-time elevation for any agent action that requires write access to sensitive systems.

5. Constrain the tool surface. In agent frameworks, explicitly allowlist the tools and API endpoints the agent may invoke — deny-by-default. Treat every MCP server as third-party code: pin versions, review tool definitions, and monitor for tool-definition drift (a known supply-chain vector in the MCP ecosystem as of early 2026).

6. Harden against credential theft from agent hosts. Enforce IMDSv2 on EC2, restrict metadata service access to required agents only via firewall rules, and alert on any process outside your cloud platform agents touching 169.254.169.254 (detections above).

7. Build the agent behavioral baseline. Deploy the KQL hunt above as a scheduled Sentinel analytic rule. An agent identity touching a new resource type for the first time should page a human. That single control would have surfaced the majority of the over-permission incidents we have investigated.

8. Plan for prompt injection as a credential abuse vector. Any agent that ingests untrusted content (emails, documents, web pages, ticket text) and also holds credentials must be treated as a confused-deputy risk. Segment these agents onto dedicated identities with the narrowest possible scopes, and consider human-in-the-loop approval gates for consequential actions (deletions, external sends, financial operations).

The organizations that get this right will not be the ones that banned AI agents — they will be the ones that gave every agent a first-class, tightly-scoped, continuously-monitored identity. Autonomy and containment are not mutually exclusive, but containment does not happen by default. Build it deliberately.

Related Resources

Security Arsenal Penetration Testing Services AlertMonitor Platform Book a SOC Assessment vulnerability-management Intel Hub

Is your security operations ready?

Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.