Back to Intelligence

Defending Against LLM Safety Classifier Bypass: Detection and Hardening Guide for AI Gateways

SA
Security Arsenal Team
October 6, 2026
10 min read

CrowdStrike recently published research detailing a deceptively simple evasion technique against large language model (LLM) safety classifiers: request fragmentation and aggregation. Instead of submitting one overtly malicious prompt — which any competent guardrail will block — the attacker decomposes the malicious objective into a series of individually benign sub-requests. Each fragment passes the safety classifier on its own merits. The attacker then aggregates the responses client-side and reconstructs the harmful output.

If your organization has deployed LLMs — whether customer-facing chatbots, internal copilots, or AI-assisted development tooling — and your trust boundary rests on a per-request safety classifier, you have a structural blind spot. This is not a theoretical jailbreak exercise. It is a repeatable, low-skill evasion pattern that works against classifiers precisely because they were designed to evaluate prompts in isolation. Security teams that have spent the last two years racing to deploy AI features are now discovering that the guardrail model they inherited does not survive contact with an adversary who thinks in sessions, not single requests.

This post breaks down how the technique works, why existing controls miss it, and — most importantly — what you can actually detect and enforce today.

Technical Analysis: Request, Aggregate, Bypass

The Attack Chain

The technique CrowdStrike describes follows a consistent three-phase pattern:

Phase 1 — Decomposition. The attacker takes a request that would trigger a safety classifier (e.g., instructions for synthesizing a dangerous substance, generating functional malware, or extracting sensitive data) and splits it into atomic, context-free sub-questions. Each fragment is engineered to be benign in isolation: "What chemical compounds are commonly found in household cleaning products?" rather than "How do I combine these into a weapon?"

Phase 2 — Distribution. The fragments are submitted across multiple requests. Sophisticated variants distribute them across sessions, API keys, accounts, or even multiple model providers to defeat session-aware analysis. Because safety classifiers typically operate as a stateless filter — evaluate prompt, return allow/deny — each request is judged without the context of its siblings.

Phase 3 — Aggregation. The attacker collects the individually benign responses and reassembles them into the complete harmful artifact. Critically, this phase happens entirely client-side. Your model never sees the malicious whole — it only ever served harmless-looking parts.

Why Existing Controls Fail

The root cause is architectural, not a tuning problem:

  • Stateless classifiers. Most deployed guardrails — including managed content-filtering services bolted onto LLM APIs — evaluate each request/response pair independently. They have no memory of what the same identity asked thirty seconds ago.
  • Semantic dilution. Fragmented prompts deliberately avoid the token patterns, keywords, and semantic signatures that classifiers are trained on. Each fragment sits comfortably inside the "benign" decision boundary.
  • Client-side aggregation is invisible. No amount of prompt filtering detects reassembly that happens on the attacker's machine. The only observable signal is the pattern of requests themselves.
  • Identity dispersion. Rotating sessions, API keys, and source IPs defeats naive per-session correlation unless the defender aggregates on stronger identity signals (authenticated user, tenant, behavioral fingerprint).

Exploitation Status

This is a demonstrated research technique published by CrowdStrike, not a vendor-specific vulnerability — there is no CVE, no patch, and no single product to update. The weakness is a class of design flaw affecting any LLM deployment whose safety enforcement is per-request and stateless. That includes a large percentage of production deployments today: direct API integrations, gateway-proxied copilots, and RAG pipelines where the guardrail sits only on the ingress prompt. Treat this as a technique in the adversarial ML playbook (MITRE ATLAS territory — evasion of ML model controls) that will be incorporated into automated jailbreak tooling, if it hasn't been already.

Detection & Response

Because the evasion lives in the request pattern, detection engineering must move up from single-event logic to behavioral aggregation. The detections below assume your LLM traffic flows through a proxy, API gateway, or managed endpoint (Azure OpenAI, Amazon Bedrock, an internal AI gateway) whose logs you are shipping to your SIEM. If you cannot currently log full LLM request metadata — source identity, timestamps, token counts, endpoint, and prompt hashes — that is your first remediation item.

Sigma Rules

The following rules target the two most reliable observables: (1) anomalous burst volume against LLM endpoints from a single source — the fingerprint of automated fragmentation frameworks — and (2) prompt content showing decomposition markers and obfuscation commonly used to smuggle fragments past classifiers. Both assume proxy/gateway log ingestion.

YAML
---
title: Suspicious Burst of Requests to LLM API Endpoints
description: Detects high-frequency request bursts to LLM API endpoints from a single source, consistent with automated prompt-fragmentation tooling used to evade stateless safety classifiers.
references:
  - https://www.crowdstrike.com/en-us/blog/how-attackers-can-bypass-llm-safety-classifiers/
  - https://atlas.mitre.org/techniques/AML.T0051/
author: Security Arsenal
id: 3f8a1b62-7c4d-4e19-9a55-2b6d8e0f1a34
date: 2026/04/06
status: experimental
logsource:
  category: proxy
detection:
  selection:
    c-uri|contains:
      - '/v1/chat/completions'
      - '/v1/completions'
      - '/v1/messages'
      - '/openai/deployments/'
      - '/generateContent'
  condition: selection
falsepositives:
  - Legitimate high-volume applications (batch summarization, code assistants). Baseline per-application request rates and alert on deviation rather than absolute counts; tune with per-source thresholding in your SIEM correlation layer.
level: medium
---
title: LLM Prompt Fragmentation and Obfuscation Markers in Request Body
description: Detects prompt content containing decomposition instructions, encoding markers, or fragmentation language commonly used to split malicious objectives into benign-appearing sub-requests that evade LLM safety classifiers.
references:
  - https://www.crowdstrike.com/en-us/blog/how-attackers-can-bypass-llm-safety-classifiers/
  - https://atlas.mitre.org/techniques/AML.T0051/
author: Security Arsenal
id: 91c4d7e0-3a2b-4f58-b671-8d3e5c90a2f7
date: 2026/04/06
status: experimental
logsource:
  category: proxy
detection:
  selection_fragmentation:
    cs-request-body|contains:
      - 'answer each part separately'
      - 'ignore the previous context'
      - 'treat each question independently'
      - 'do not consider earlier messages'
      - 'decode the following base64'
      - 'combine the answers'
  selection_endpoint:
    c-uri|contains:
      - '/chat/completions'
      - '/completions'
      - '/messages'
  condition: selection_fragmentation and selection_endpoint
falsepositives:
  - Legitimate prompt-engineering templates used by internal developers. Maintain an allowlist of known internal application identities and suppress by authenticated service principal.
level: medium

A candid note on rule quality: per-request content matching will always lag attacker creativity, and a raw rate rule without per-source baselining will drown you. The durable detection is session-aware aggregation in your SIEM, which is where the KQL below earns its keep.

KQL Hunt — Microsoft Sentinel

This query hunts for the fragmentation pattern directly: a single identity issuing an abnormally high number of distinct, low-complexity prompts to LLM endpoints within a short window — the statistical shape of decomposition. It assumes LLM gateway or proxy logs are ingested into CommonSecurityLog (CEF) or a custom table; adjust table and field names to your ingestion pipeline.

KQL — Microsoft Sentinel / Defender
// Hunt: LLM safety-classifier evasion via request fragmentation
// Detects identities issuing high volumes of distinct, short prompts to LLM endpoints in a tight window
let Lookback = 24h;
let BurstWindow = 15m;
let DistinctPromptThreshold = 20;
let LLMEndpoints = dynamic(["/v1/chat/completions", "/v1/messages", "/openai/deployments/", "/generateContent", "/v1/completions"]);
CommonSecurityLog
| where TimeGenerated > ago(Lookback)
| where RequestURL has_any (LLMEndpoints)
| extend SourceIdentity = coalesce(Extention, SourceUserID, SourceIP)
| summarize
    RequestCount = count(),
    DistinctPrompts = dcount(AdditionalExtensions),
    FirstRequest = min(TimeGenerated),
    LastRequest = max(TimeGenerated),
    DistinctSessions = dcount(DeviceCustomString1),
    Endpoints = make_set(RequestURL)
    by SourceIdentity, SourceIP, bin(TimeGenerated, BurstWindow)
| where DistinctPrompts >= DistinctPromptThreshold
| extend SessionDispersionRatio = round(todouble(DistinctSessions) / todouble(RequestCount), 2)
| project TimeGenerated, SourceIdentity, SourceIP, RequestCount, DistinctPrompts, DistinctSessions, SessionDispersionRatio, FirstRequest, LastRequest, Endpoints
| order by DistinctPrompts desc

Tune DistinctPromptThreshold against your own baselines — a code-assistant workload and a customer chatbot have very different normal shapes. The SessionDispersionRatio column surfaces the deliberate session-rotation variant of the technique: legitimate users cluster in one session; fragmentation frameworks scatter across many.

Velociraptor VQL — Shadow AI Discovery

One underappreciated defensive angle: fragmentation attacks are far easier to run against unmonitored LLM access. Employees or attackers standing up local model clients (Ollama, LM Studio, text-generation-webui) create AI capability entirely outside your gateway and logging fabric. This VQL artifact hunts endpoints for running local LLM clients and their model caches — your shadow-AI inventory is part of your classifier-evasion exposure.

VQL — Velociraptor
-- Hunt for unauthorized local LLM clients (shadow AI) on endpoints
-- Covers common runtimes used to run models outside monitored gateways
SELECT Pid, Name, Exe, CommandLine, Username, CreateTime
FROM pslist()
WHERE Name =~ '(?i)ollama|lmstudio|lm-studio|text-generation-webui|llamafile|koboldcpp|gpt4all|localai'
   OR CommandLine =~ '(?i)ollama (serve|run)|lmstudio|llama.cpp|server.*--model'

Pair this with a glob() sweep over default model-cache directories (C:/Users/*/.ollama/models/**, C:/Users/*/.cache/lm-studio/**) for a complete picture of unsanctioned model deployments.

Hardening Script

For organizations on Azure OpenAI, the highest-leverage controls are (a) ensuring content-filter configuration is actually enforced and (b) turning on diagnostic logging so the KQL above has data to chew on. This PowerShell script audits every Azure OpenAI account in a subscription for diagnostic settings and reports accounts where prompt/response logging is not flowing to a Log Analytics workspace.

PowerShell
# Audit Azure OpenAI accounts for missing diagnostic logging (defense against classifier-evasion blind spots)
# Requires: Az module, Reader on subscription
$subscriptions = Get-AzSubscription
foreach ($sub in $subscriptions) {
    Set-AzContext -SubscriptionId $sub.Id | Out-Null
    $accounts = Get-AzResource -ResourceType "Microsoft.CognitiveServices/accounts" |
        Where-Object { $_.Kind -eq "OpenAI" }
    foreach ($acct in $accounts) {
        $diag = Get-AzDiagnosticSetting -ResourceId $acct.ResourceId -ErrorAction SilentlyContinue
        $hasLogging = $diag | Where-Object {
            $_.WorkspaceId -and ($_.Log | Where-Object { $_.Enabled -and ($_.Category -in @("AuditEvent","RequestResponse") -or $_.CategoryGroup -eq "allLogs") })
        }
        [PSCustomObject]@{
            Subscription   = $sub.Name
            Account        = $acct.Name
            ResourceGroup  = $acct.ResourceGroupName
            LoggingEnabled = [bool]$hasLogging
        }
    }
}
# Remediate an unprotected account (example):
# Set-AzDiagnosticSetting -ResourceId $acct.ResourceId -WorkspaceId "<LAW-resource-id>" -Enabled $true -Category "AuditEvent","RequestResponse"

Remediation

There is no patch for a design flaw — there is architecture. Prioritize in this order:

  1. Move safety evaluation from per-request to per-session. Your guardrail must evaluate the cumulative semantic content of a conversation, not each prompt in isolation. If your managed content-filtering service is stateless, place a session-aware aggregation layer in front of it (most enterprise AI gateways now support conversation-context evaluation — verify yours does and that it is enabled, not just licensed).
  2. Enforce strong, non-rotatable identity on all LLM traffic. Authenticate every request to a named user or service principal. Block anonymous/shared-key access. Fragmentation across identities only works if identities are cheap.
  3. Log everything, centrally. Full request metadata (identity, timestamp, endpoint, token counts, prompt hash) into your SIEM. Without this, every detection above is theoretical. Run the audit script this week.
  4. Rate-limit and anomaly-alert at the gateway. Establish per-identity baselines and alert on burst deviations and high distinct-prompt counts. This is cheap and catches the automated tooling that makes fragmentation practical at scale.
  5. Add output-side aggregation checks. Classify not just each response but the rolling combination of recent responses per session. Harmful reassembly leaves a detectable trail when consecutive benign outputs form a sensitive composite.
  6. Inventory and eliminate shadow AI. Local model clients bypass every control you deploy. Use the VQL artifact above, and give users a sanctioned, monitored alternative so the incentive to go around you disappears.
  7. Red-team your own guardrails. Add fragmentation-and-aggregation scenarios to your AI red-team scope. If your penetration testing program isn't testing your LLM trust boundaries, it has a gap the size of your AI footprint.

Related Resources

Security Arsenal Penetration Testing Services AlertMonitor Platform Book a SOC Assessment vulnerability-management Intel Hub

Is your security operations ready?

Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.