Back to Intelligence

PoeLLM Malware Infects Exposed AI Servers for Cryptomining: Detection and Hardening Guide

SA
Security Arsenal Team
October 8, 2026
11 min read

Security researchers have disclosed an active cryptomining campaign deploying malware tracked as PoeLLM against internet-exposed AI servers. The operation doesn't stop at hijacking GPU and CPU cycles for illicit mining — compromised hosts are being weaponized as network scanners and launchpads for follow-on attacks against other exposed services. If your organization runs self-hosted LLM inference — Ollama, vLLM, or any OpenAI-compatible API endpoint — on a cloud VM with a public IP, you are in the blast radius.

This is the pattern we've been warning clients about since AI adoption went vertical in 2024 and 2025: engineering teams spin up inference endpoints for demos and internal tooling, skip authentication, bind to 0.0.0.0, and leave the box unmanaged by EDR, patch management, or the SIEM. These servers ship with high-end GPUs, generous bandwidth, and frequently over-privileged cloud IAM roles. To a cryptomining operator, that's a better target than a thousand compromised WordPress sites. The fact that PoeLLM also converts victims into scanning infrastructure means a single exposed inference server can become the source of attacks against your customers and partners — creating legal and reputational exposure well beyond an inflated cloud bill.

Technical Analysis: How the PoeLLM Campaign Operates

Target Profile

The campaign hunts for AI services exposed directly to the internet without authentication. Typical victim characteristics:

  • Self-hosted LLM inference servers (Ollama on TCP/11434, vLLM on TCP/8000, text-generation-webui, LocalAI, and similar OpenAI-compatible APIs) bound to all interfaces
  • Cloud VMs (AWS, GCP, Azure, and low-cost GPU VPS providers) deployed outside corporate landing zones — often "shadow AI" stood up by developers or data science teams
  • Hosts running as root or with sudo-capable service accounts, frequently without endpoint detection coverage
  • Containers with the Docker socket mounted or hosts with exposed Docker daemon APIs (TCP/2375)

Attack Chain

From a defender's perspective, the campaign follows a repeatable sequence:

  1. Discovery/reconnaissance — The operators (and previously compromised PoeLLM victims acting as scanners) sweep the internet for common AI service ports and API fingerprints (/api/tags on Ollama, /v1/models on OpenAI-compatible endpoints).
  2. Initial access — Unauthenticated API access is abused to execute code or pull malicious models/artifacts onto the host. Where the service itself doesn't permit direct command execution, attackers chain misconfigurations: exposed Docker APIs, co-located Jupyter instances, or default credentials on adjacent management interfaces.
  3. Payload staging — A loader script is fetched from attacker infrastructure via curl or wget piped directly to a shell, written to a world-writable path such as /tmp or /dev/shm to evade naïve file integrity monitoring.
  4. Miner deployment — The PoeLLM payload deploys a cryptominer (XMRig or a derivative), typically masquerading under a legitimate-looking process name and tuning itself to avoid obvious CPU saturation — on GPU hosts, mining workloads contend directly with inference jobs.
  5. Persistence — Cron entries, systemd service units, and shell profile modifications (~/.bashrc, /etc/profile.d/) are used to survive reboots. Some variants kill competing miners and disable cloud monitoring agents to reduce the chance of detection.
  6. Weaponization — The host is enrolled into the campaign's scanning infrastructure and begins probing the internet for additional exposed AI services, effectively making each victim a force multiplier and laundering the scanner's origin through legitimate cloud IPs.

Exploitation Status

This is confirmed, active, in-the-wild exploitation of exposed services — not a theoretical threat. No CVE is associated with the campaign; the vulnerability is architectural: unauthenticated, internet-facing AI inference endpoints. There is no vendor patch to wait for. The fix is entirely in your deployment posture.

Detection & Response

The detections below target the highest-fidelity, lowest-noise behaviors in this chain: shell-piped downloads from AI service processes, miner execution from ephemeral paths, cron/systemd persistence, and outbound scanning behavior from hosts that have no business initiating wide internet connections.

YAML
---
title: AI Inference Service Spawning Shell or Download Utility
id: 3f9c2a71-8b44-4e1a-9c6d-2a7f5b8e1d03
status: experimental
description: Detects LLM inference server processes (Ollama, vLLM, text-generation-webui, LocalAI) spawning shells or download utilities, consistent with PoeLLM initial-access and payload staging behavior.
references:
  - https://www.bleepingcomputer.com/news/security/poellm-malware-infects-exposed-ai-servers-in-cryptomining-attacks/
  - https://attack.mitre.org/techniques/T1190/
author: Security Arsenal
date: 2026/04/10
tags:
  - attack.initial_access
  - attack.t1190
  - attack.execution
  - attack.t1059.004
logsource:
  category: process_creation
  product: linux
detection:
  selection_parent:
    ParentImage|contains:
      - '/ollama'
      - '/vllm'
      - '/text-generation-webui'
      - '/localai'
      - 'python'
  selection_child:
    Image|endswith:
      - '/sh'
      - '/bash'
      - '/curl'
      - '/wget'
      - '/chmod'
      - '/crontab'
  condition: selection_parent and selection_child
falsepositives:
  - Legitimate model download tooling or custom inference wrappers that shell out to curl
level: high
---
title: Cryptominer Execution from Ephemeral or World-Writable Path
id: 8e1d4b62-5c09-4f3a-b712-9d4e6a2c7f15
status: experimental
description: Detects execution of binaries from /tmp, /dev/shm, or /var/tmp combined with known miner names or mining pool command-line arguments, as used by the PoeLLM payload staging stage.
references:
  - https://www.bleepingcomputer.com/news/security/poellm-malware-infects-exposed-ai-servers-in-cryptomining-attacks/
  - https://attack.mitre.org/techniques/T1496/
author: Security Arsenal
date: 2026/04/10
tags:
  - attack.impact
  - attack.t1496
  - attack.defense_evasion
  - attack.t1036
logsource:
  category: process_creation
  product: linux
detection:
  selection_path:
    Image|contains:
      - '/tmp/'
      - '/dev/shm/'
      - '/var/tmp/'
  selection_miner:
    CommandLine|contains:
      - 'xmrig'
      - '--donate-level'
      - 'stratum+tcp'
      - 'stratum+ssl'
      - 'moneroocean'
      - 'supportxmr'
      - '--coin'
      - '-o pool.'
  selection_name:
    Image|endswith:
      - '/xmrig'
      - '/kdevtmpfsi'
      - '/kinsing'
      - '/pnscan'
      - '/masscan'
  condition: selection_name or (selection_path and selection_miner)
falsepositives:
  - Rare; ephemeral paths combined with mining-pool arguments are almost never legitimate
level: critical
---
title: Persistence via Cron or Systemd from User-Writable Locations
id: 2c7a5e94-6d31-4b8f-a023-5e9b1c4d8f62
status: experimental
description: Detects cron or systemd persistence referencing scripts or binaries in /tmp, /dev/shm, or /var/tmp — a common PoeLLM persistence mechanism for surviving reboots on compromised AI servers.
references:
  - https://www.bleepingcomputer.com/news/security/poellm-malware-infects-exposed-ai-servers-in-cryptomining-attacks/
  - https://attack.mitre.org/techniques/T1053/003/
  - https://attack.mitre.org/techniques/T1543/002/
author: Security Arsenal
date: 2026/04/10
tags:
  - attack.persistence
  - attack.t1053.003
  - attack.t1543.002
logsource:
  category: file_event
  product: linux
detection:
  selection_dir:
    TargetFilename|startswith:
      - '/etc/cron'
      - '/var/spool/cron'
      - '/etc/systemd/system/'
      - '/usr/lib/systemd/system/'
  selection_content:
    TargetFilename|contains:
      - 'tmp'
      - 'shm'
  selection_generic:
    TargetFilename|startswith:
      - '/etc/cron.d/'
      - '/var/spool/cron/crontabs/'
  condition: (selection_dir and selection_content) or selection_generic
falsepositives:
  - Legitimate cron entries created by configuration management (Ansible, Chef) — baseline and exclude known provisioning sources
level: high
KQL — Microsoft Sentinel / Defender
// PoeLLM hunt: AI servers fetching payloads, miner behavior, and outbound scanning
// Assumes Linux Syslog/auditd ingestion into Sentinel (Syslog table) plus Defender for Endpoint where deployed.

// 1. Shell-piped downloads and miner indicators on Linux hosts via auditd/Syslog
Syslog
| where TimeGenerated > ago(7d)
| where ProcessName in~ ("bash", "sh", "curl", "wget", "crontab", "systemctl")
| where SyslogMessage has_any (
    "curl -s", "wget -q", "| sh", "| bash", "|sh", "|bash",
    "xmrig", "stratum+tcp", "stratum+ssl", "--donate-level",
    "/dev/shm", "/var/tmp/", "moneroocean", "supportxmr")
| extend SuspiciousPath = iif(SyslogMessage has_any ("/tmp/", "/dev/shm/", "/var/tmp/"), true, false)
| where SuspiciousPath or SyslogMessage has_any ("xmrig", "stratum+", "--donate-level")
| summarize FirstSeen=min(TimeGenerated), LastSeen=max(TimeGenerated), SampleCommands=make_set(SyslogMessage, 10)
  by Computer, ProcessName
| order by FirstSeen desc;

// 2. Outbound scanning behavior: single host fanning out to many distinct public IPs on AI service ports
// (Ollama 11434, vLLM/OpenAI-compatible 8000/8080, Docker API 2375) — the PoeLLM weaponization stage.
Syslog
| where TimeGenerated > ago(24h)
| where SyslogMessage has_any ("DPT=11434", "DPT=8000", "DPT=8080", "DPT=2375") and SyslogMessage has "SYN"
| parse SyslogMessage with * "SRC=" SourceIP " " * "DST=" DestIP " " * "DPT=" DestPort " " *
| where not(ipv4_is_private(DestIP))
| summarize DistinctTargets=dcount(DestIP), Ports=make_set(DestPort) by Computer, SourceIP
| where DistinctTargets > 50
| order by DistinctTargets desc;
VQL — Velociraptor
-- Velociraptor hunt: identify PoeLLM-style compromise on Linux AI servers.
-- Collects suspicious processes, cron/systemd persistence, and ephemeral-path binaries.
SELECT * FROM chain(
  a = {
    -- Stage 1: live processes executing from ephemeral paths or with miner strings
    SELECT Pid, Name, Exe, Commandline, Username
    FROM pslist()
    WHERE Exe =~ '/(tmp|dev/shm|var/tmp)/'
       OR Commandline =~ '(?i)(xmrig|stratum\+|--donate-level|moneroocean|supportxmr)'
  },
  b = {
    -- Stage 2: cron persistence artifacts
    SELECT FullPath AS Artifact, 'cron' AS Type, Size, Mtime
    FROM glob(globs=['/etc/cron.d/*', '/etc/cron.daily/*', '/var/spool/cron/*', '/var/spool/cron/crontabs/*'])
  },
  c = {
    -- Stage 3: recently modified or non-standard systemd units
    SELECT FullPath AS Artifact, 'systemd' AS Type, Size, Mtime
    FROM glob(globs=['/etc/systemd/system/*.service'])
    WHERE Mtime > now() - 604800  -- units modified in the last 7 days
  },
  d = {
    -- Stage 4: binaries dropped in world-writable/ephemeral paths
    SELECT FullPath AS Artifact, 'dropped_binary' AS Type, Size, Mtime
    FROM glob(globs=['/tmp/*', '/dev/shm/*', '/var/tmp/*'])
    WHERE NOT IsDir
  }
)
Bash / Shell
#!/bin/bash
# poeLLM_triage_harden.sh — triage and harden exposed AI inference servers.
# Run as root. Review output before enabling the firewall section.
set -euo pipefail

echo "=== [1] Suspicious processes: miners, scanners, ephemeral-path binaries ==="
ps auxww | grep -Ei 'xmrig|kinsing|kdevtmpfsi|masscan|pnscan|stratum|donate-level' | grep -v grep || echo "None found"
ls -la /tmp /dev/shm /var/tmp 2>/dev/null | grep -vE '^total|^d' || true

echo "=== [2] Persistence: cron and systemd ==="
crontab -l 2>/dev/null; ls -la /etc/cron.d/ /var/spool/cron/ 2>/dev/null
grep -rEl '/tmp/|/dev/shm/|/var/tmp/' /etc/systemd/system/ 2>/dev/null || echo "No systemd units referencing ephemeral paths"
systemctl list-unit-files --state=enabled | grep -viE 'ssh|cron|systemd|network|rsyslog|cloud-init|docker|containerd' || true

echo "=== [3] Exposure check: is an AI service listening on a public interface? ==="
ss -tlnp | grep -E ':(11434|8000|8080|2375|5000)\b' || echo "No common AI service ports listening"

echo "=== [4] Outbound scanning connections from this host ==="
ss -tnp state established | awk '{print $5}' | cut -d: -f1 | sort | uniq -c | sort -rn | head -20

echo "=== [5] HARDEN: bind Ollama to localhost only (systemd override) ==="
if systemctl list-unit-files | grep -q '^ollama.service'; then
  mkdir -p /etc/systemd/system/ollama.service.d
  cat > /etc/systemd/system/ollama.service.d/override.conf <<'EOF'
[Service]
Environment="OLLAMA_HOST=127.0.0.1:11434"
EOF
  systemctl daemon-reload
  systemctl restart ollama
  echo "Ollama bound to 127.0.0.1 — expose it via an authenticated reverse proxy if remote access is required."
fi

echo "=== [6] HARDEN: default-deny egress for mining pools and scanning (ufw) ==="
# Uncomment after validating dependencies — blocks outbound stratum and AI-port fan-out.
# ufw default deny outgoing
# ufw allow out 53; ufw allow out 80; ufw allow out 443; ufw allow out 123/udp
# ufw enable

echo "=== Done. Re-image any host with confirmed miner artifacts — do not trust in-place cleanup. ==="

Remediation and Hardening

There is no patch because there is no product vulnerability — the remediation is architectural and operational. Prioritize in this order:

  1. Inventory your exposure today. Enumerate every public IP in your cloud estates and scan for listening AI service ports: Ollama (11434), vLLM and OpenAI-compatible APIs (8000, 8080), text-generation-webui (7860, 5000), Docker API (2375/2376), and Jupyter (8888). Shodan/Censys queries against your own ASN take ten minutes and regularly surface shadow AI infrastructure your CMDB doesn't know about. Do this before the PoeLLM scanners do it again.
  2. Kill unauthenticated exposure. AI inference endpoints must never be reachable from the internet without an identity-aware proxy, VPN, or mTLS in front. Bind services to localhost or a private interface and front them with an authenticated reverse proxy (nginx with OIDC, Cloudflare Access, or equivalent). Authentication on the inference API itself is a secondary control, not the primary one.
  3. Contain and re-image confirmed victims. A host with PoeLLM artifacts is a host with unknown persistence depth and unknown secondary tooling — it was also, by definition, used to attack third parties from your IP space. Isolate at the security-group/NACL layer, snapshot disks for forensics, rotate every credential and cloud token that host could access (instance profile credentials are the big one — assume they were exfiltrated), and rebuild from a known-good image. Do not attempt in-place cleanup.
  4. Deploy endpoint telemetry to AI workloads. These servers are frequently exempted from EDR "because GPU performance." That exemption is exactly why miners thrive there. A lightweight eBPF-based agent or at minimum auditd with process-execution logging, shipped to your SIEM, closes the gap the detections above depend on.
  5. Set GPU and egress alerting. Cryptominers have an unmissable signature if you're looking: sustained GPU utilization outside inference jobs, and outbound connections to known mining pool domains/IPs. Alert on both. Default-deny egress from inference subnets — these hosts need model registries and package mirrors, not the open internet.
  6. Govern shadow AI. The root cause is almost always an ungoverned deployment. Establish an approval and registration path for self-hosted inference so security knows these assets exist, and enforce baseline hardening (non-root service accounts, patched base images, no Docker socket exposure) through your pipeline, not through hope.

If you find evidence of PoeLLM or a comparable compromise in your environment and need support with scoping, credential rotation, or forensic analysis of the weaponization stage, engage your IR retainer immediately — the outbound scanning activity means third-party notification obligations may already be accruing.

Related Resources

Security Arsenal Incident Response Services AlertMonitor Platform Book a SOC Assessment incident-response Intel Hub

Is your security operations ready?

Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.