Back to Intelligence

Suspicious User-Agent Strings in Web Logs: Honeypot Lessons for Detection and Threat Hunting

SA
Security Arsenal Team
October 4, 2026
10 min read

A recent SANS Internet Storm Center diary entry (October 4th) highlighted something every veteran SOC analyst recognizes: reviewing User-Agent strings in honeypot logs is alternately amusing and deeply informative. Buried in that traffic are scanner fingerprints, botnet identifiers, copy-paste exploit kits, and the occasional attacker who forgot to change a default header. The defensive lesson is not the humor — it is that User-Agent analysis remains one of the cheapest, highest-signal detection surfaces available to defenders, and most organizations barely look at it.

This matters now more than ever. Internet-wide scanning is constant, automated exploitation follows public PoC releases within hours, and a large share of that traffic announces itself — knowingly or not — through the User-Agent header. If your SOC is not hunting on this field in 2026, you are leaving free detections on the table.

Technical Analysis: What User-Agent Strings Reveal

Why honeypots see what you don't

Honeypots like the DShield/ISC sensor network have no legitimate users. Every request is, by definition, reconnaissance, scanning, or attack traffic. That makes them a perfect laboratory for observing attacker tooling. Your production web servers are noisier, but the same tooling shows up in your logs — you just have to separate it from legitimate browser and API traffic.

Common categories of suspicious User-Agent strings

From honeypot and production log analysis, suspicious UAs generally fall into these buckets:

  1. Offensive security tool defaults — tools that advertise themselves unless configured otherwise: sqlmap/1.x, Nikto, Nuclei, Nmap Scripting Engine, masscan, zgrab, gobuster, WPScan, Acunetix, hydra. Anyone running these against your perimeter without a signed rules-of-engagement document is hostile until proven otherwise.
  2. Automation library defaults — python-requests/x.y, Go-http-client/1.1, curl/x.y, libwww-perl, Java/1.x, Wget. These are not inherently malicious, but a python-requests UA hammering /wp-login.php, /admin, or .env paths is a strong indicator of scripted attack activity.
  3. Empty or malformed UAs — a missing User-Agent header, a single dash, or garbage strings. Legitimate browsers always send a UA. Empty UAs on web requests correlate strongly with crude bots and exploit scripts.
  4. Impossible or anachronistic combinations — ancient browser versions (e.g., MSIE 6.0) claiming to originate from mobile devices, or UAs that don't match the TLS client fingerprint. Sophisticated bots rotate through UA lists; mismatches between the claimed client and observed behavior (TLS cipher suites, header ordering, request cadence) expose them.
  5. Injection and probe payloads in the UA field itself — attackers routinely place shell commands, template injection strings (e.g., ${jndi:ldap://...} style payloads), and XSS probes in the User-Agent header, betting that some backend system will log, parse, or render it unsafely. Any UA containing ${, $(, backticks, or <script> should be treated as an attack attempt against your log-processing pipeline, not just your web app.

Exploitation status and threat landscape

This is not a single-CVE story — it is the ambient background radiation of the internet. Automated scanning and exploitation traffic is continuous and confirmed active against every publicly routable IP. The practical implication: any detectable-by-UA tool hitting your perimeter is already in-the-wild, by definition. Your exposure depends on whether the underlying target (an unpatched edge device, a forgotten WordPress install, an exposed admin panel) is reachable.

The spoofing caveat — read this before deploying anything

User-Agent is attacker-controlled and trivially spoofed. Never use UA matching as a blocking-only control in production, and never treat a legitimate-looking UA as proof of legitimacy. UA-based detections are most powerful as:

  • Alerting and hunting signals (default tool UAs are noisy attackers or lazy attackers — both worth catching)
  • Correlation enrichers (UA anomaly + sensitive path + high request rate = escalation)
  • Deception triggers (route known scanner UAs to honeypot endpoints for high-fidelity alerting)

Detection & Response

Sigma Rules

YAML
---
title: Offensive Security Tool User-Agent Detected in Web Request
id: 3f8a2c71-9b4e-4d1a-a6c2-7e5f0d8b9c1a
status: experimental
description: Detects web requests carrying default User-Agent strings of known offensive security and scanning tools (sqlmap, Nuclei, Nikto, masscan, etc.). These defaults indicate untuned or automated attack tooling against the web perimeter.
references:
  - https://isc.sans.edu/diary/rss/33394
  - https://attack.mitre.org/techniques/T1595/
author: Security Arsenal
date: 2026/01/12
tags:
  - attack.reconnaissance
  - attack.t1595
  - attack.t1595.002
logsource:
  category: webserver
detection:
  selection:
    c-useragent|contains:
      - 'sqlmap'
      - 'Nuclei'
      - 'Nikto'
      - 'masscan'
      - 'zgrab'
      - 'gobuster'
      - 'dirbuster'
      - 'WPScan'
      - 'Acunetix'
      - 'hydra'
      - 'Nmap Scripting Engine'
      - 'nferno'
  condition: selection
falsepositives:
  - Authorized penetration tests and vulnerability scanners (whitelist by source IP during sanctioned test windows)
  - Researchers and search engine security crawlers
level: high
---
title: Automation Library User-Agent Accessing Sensitive Web Paths
id: 8c1d4e62-2a7f-4b9c-b3d5-1f6a9e0c2d4b
status: experimental
description: Detects scripting library User-Agents (python-requests, Go-http-client, curl, wget, libwww-perl) requesting sensitive application paths such as admin panels, environment files, or WordPress login pages. Correlating a non-browser UA with a sensitive target path sharply reduces noise.
references:
  - https://isc.sans.edu/diary/rss/33394
  - https://attack.mitre.org/techniques/T1595/
author: Security Arsenal
date: 2026/01/12
tags:
  - attack.reconnaissance
  - attack.t1595
  - attack.t1190
logsource:
  category: webserver
detection:
  selection_ua:
    c-useragent|contains:
      - 'python-requests'
      - 'Go-http-client'
      - 'curl/'
      - 'Wget/'
      - 'libwww-perl'
      - 'Java/1.'
      - 'okhttp'
  selection_path:
    cs-uri-stem|contains:
      - '/wp-login'
      - '/wp-admin'
      - '/admin'
      - '/.env'
      - '/.git'
      - '/phpmyadmin'
      - '/xmlrpc.php'
      - '/actuator'
      - '/console'
  condition: selection_ua and selection_path
falsepositives:
  - Legitimate API clients and health-check agents (tune path list to your application inventory)
  - Internal monitoring tooling hitting admin endpoints
level: medium
---
title: Suspicious Payload Content in User-Agent Header
id: b5e7f0a3-6d2c-48e1-9a4b-3c7d1e5f8a0b
status: experimental
description: Detects User-Agent headers containing shell metacharacters, template injection markers, or script tags, indicating an attempt to attack log-processing, analytics, or backend systems that consume the UA field.
references:
  - https://isc.sans.edu/diary/rss/33394
  - https://attack.mitre.org/techniques/T1190/
author: Security Arsenal
date: 2026/01/12
tags:
  - attack.initial_access
  - attack.t1190
logsource:
  category: webserver
detection:
  selection:
    c-useragent|contains:
      - '${'
      - '$('
      - '`'
      - '<script'
      - '{{'
      - '|'
      - ';'
  filter_browsers:
    c-useragent|contains:
      - 'Mozilla/'
      - 'AppleWebKit'
      - 'Chrome/'
      - 'Safari/'
  condition: selection and not filter_browsers
falsepositives:
  - Rare malformed legitimate clients; validate UA entropy before escalation
level: high

KQL — Microsoft Sentinel (CEF/Syslog/WAF ingestion)

KQL — Microsoft Sentinel / Defender
// Hunt for scanner and automation User-Agents across web/proxy telemetry ingested via CEF or Syslog
let ScannerUA = dynamic(["sqlmap", "Nuclei", "Nikto", "masscan", "zgrab", "gobuster", "WPScan", "Acunetix", "hydra", "Nmap Scripting Engine"]);
let AutomationUA = dynamic(["python-requests", "Go-http-client", "curl/", "Wget/", "libwww-perl", "okhttp"]);
let SensitivePaths = dynamic(["/wp-login", "/wp-admin", "/admin", "/.env", "/.git", "/phpmyadmin", "/xmlrpc.php", "/actuator"]);
union isfuzzy=true
  (CommonSecurityLog
   | extend UA = tostring(RequestClientApplication), Uri = tostring(RequestURL)),
  (Syslog
   | where SyslogMessage has_any (ScannerUA) or SyslogMessage has_any (AutomationUA)
   | extend UA = extract(@"User-Agent[":]?\s*([^\"\r\n]+)", 1, SyslogMessage), Uri = extract(@"(GET|POST|HEAD|PUT|DELETE)\s+([^\s]+)", 2, SyslogMessage))
| where isnotempty(UA)
| extend Verdict = case(
    UA has_any (ScannerUA), "Offensive Tool Default UA",
    UA has_any (AutomationUA) and Uri has_any (SensitivePaths), "Automation UA + Sensitive Path",
    UA has @"\$\{|\$\(|`|<script", "Payload in UA Header",
    isempty(UA) or UA == "-", "Empty User-Agent",
    "")
| where isnotempty(Verdict)
| summarize RequestCount = count(), DistinctURIs = dcount(Uri), SampleURIs = make_set(Uri, 10)
    by SourceIP, UA, Verdict, bin(TimeGenerated, 1h)
| order by RequestCount desc;

Tune the source-IP whitelist for your sanctioned scanner ranges (Qualys, Tenable, internal red team) before enabling alerting on this query. The hourly summarization keeps volume manageable while preserving the source/UA/verdict pivot points an analyst needs.

Velociraptor VQL — Web Server Log Hunting

For web servers in scope of an investigation, hunt access logs directly from the endpoint:

VQL — Velociraptor
-- Hunt web server access logs for scanner User-Agents and payload-in-UA indicators
LET suspicious = 'sqlmap|Nuclei|Nikto|masscan|zgrab|gobuster|WPScan|Acunetix|hydra|python-requests|Go-http-client|\\$\\{|<script'
SELECT FullPath AS LogFile,
       Line AS LogEntry,
       timestamp(string=now()) AS HuntTime
FROM foreach(
  row={
    SELECT FullPath FROM glob(globs=['/var/log/nginx/access.log*', '/var/log/apache2/access.log*', '/var/log/httpd/access_log*'])
  },
  query={
    SELECT FullPath, Line
    FROM parse_lines(filename=FullPath)
    WHERE Line =~ suspicious
  })
LIMIT 500

This gives IR responders a fast way to confirm scanner exposure on a specific host without waiting on SIEM ingestion lag — valuable when triaging a newly disclosed exploit against your stack.

Remediation / Audit Script

Bash / Shell
#!/bin/bash
# ua-audit.sh — Audit web access logs for suspicious User-Agent activity
# Run on web servers or against centralized log archives. Requires: grep, zgrep (for rotated logs)

LOG_DIRS="/var/log/nginx /var/log/apache2 /var/log/httpd"
OUT="ua-audit-$(date +%Y%m%d-%H%M).txt"
SCANNER_UA='sqlmap|Nuclei|Nikto|masscan|zgrab|gobuster|dirbuster|WPScan|Acunetix|hydra|Nmap Scripting Engine'
AUTO_UA='python-requests|Go-http-client|curl/|Wget/|libwww-perl'
PAYLOAD_UA='\$\{|\$\(|<script|\{\{'

echo "=== User-Agent Security Audit: $(date) ===" | tee "$OUT"

for DIR in $LOG_DIRS; do
  [ -d "$DIR" ] || continue
  echo "" | tee -a "$OUT"
  echo "## Scanning $DIR" | tee -a "$OUT"

  echo "-- [1] Offensive tool User-Agents (top sources):" | tee -a "$OUT"
  zgrep -hEi "$SCANNER_UA" "$DIR"/access*.log* 2>/dev/null \
    | awk '{print $1}' | sort | uniq -c | sort -rn | head -20 | tee -a "$OUT"

  echo "-- [2] Automation UAs hitting sensitive paths:" | tee -a "$OUT"
  zgrep -hEi "($AUTO_UA)" "$DIR"/access*.log* 2>/dev/null \
    | grep -Ei '/wp-login|/wp-admin|/\.env|/\.git|/phpmyadmin|/xmlrpc\.php|/actuator' \
    | awk '{print $1, $7}' | sort | uniq -c | sort -rn | head -20 | tee -a "$OUT"

  echo "-- [3] Payload content in User-Agent field:" | tee -a "$OUT"
  zgrep -hE "$PAYLOAD_UA" "$DIR"/access*.log* 2>/dev/null | head -20 | tee -a "$OUT"

  echo "-- [4] Empty or dash User-Agents:" | tee -a "$OUT"
  zgrep -hE '"(-|)"$|"-$|""' "$DIR"/access*.log* 2>/dev/null \
    | awk '{print $1}' | sort | uniq -c | sort -rn | head -20 | tee -a "$OUT"
done

echo "" | tee -a "$OUT"
echo "Audit complete. Review $OUT and feed confirmed-malicious source IPs to your blocklist/threat intel pipeline." | tee -a "$OUT"

Remediation and Hardening Recommendations

There is no vendor patch for this class of threat — the remediation is architectural and procedural:

  1. Log the User-Agent everywhere, and keep it. Verify that your WAF, load balancer, reverse proxy, and application logs all capture the full UA header (truncation at 128 characters can cut off payload-in-UA evidence). Confirm the field is parsed into a searchable field in your SIEM, not buried in a raw message blob.
  2. Deploy the detections above as alerting, not just hunting. Default offensive-tool UAs are a high-fidelity signal. Automation-UA-plus-sensitive-path is medium fidelity and needs per-environment tuning of the path list against your actual application inventory.
  3. Block known-bad at the edge, alert on everything else. WAF and CDN rules (Cloudflare, AWS WAF, ModSecurity with the OWASP Core Rule Set scanner-detection rules) can drop default scanner UAs before they touch your origin. Do not rely on UA blocking as your primary control — it stops lazy attackers only.
  4. Rate-limit and tarpit automation UAs on sensitive paths. Login pages, password reset, and admin panels should have strict rate limits regardless of UA. A python-requests client attempting 500 logins per minute should trip account lockout and IP throttling long before your SOC reads an alert.
  5. Stand up deception. Route requests with known scanner UAs to a honeypot vhost or canary endpoint. Any interaction there is hostile by definition — this converts a noisy signal into a near-zero-false-positive tripwire.
  6. Sanitize UA fields in your log pipeline. Because attackers place injection payloads in the UA header, ensure your log management platform, ticketing integrations, and any dashboards that render UA strings treat the field as untrusted input. Payload-in-UA detection exists precisely because this field gets rendered somewhere.
  7. Baseline your legitimate UA population. Build an allowlist of UAs your real users and sanctioned scanners produce. Deviations — new automation UAs from internal subnets, for instance — can indicate compromised internal hosts running scanning tools, turning an external-recon detection into an internal-compromise detection.
  8. Feed findings to threat intel. Confirmed-malicious source IPs and UA patterns from your logs should flow into your blocklists, your MDR provider's analytics, and (where appropriate) community sharing like DShield itself — the same ecosystem that produced this diary entry.

The ISC honeypot observations are a reminder that attackers are constantly probing, and many of them sign their work in the User-Agent field. The defenders who win are the ones actually reading the signature.

Related Resources

Security Arsenal Penetration Testing Services AlertMonitor Platform Book a SOC Assessment vulnerability-management Intel Hub

Is your security operations ready?

Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.