Back to Intelligence

CVE-2026-103040: LightLLM Profiler Service Unauthenticated RCE (CVSS 9.8) — Detection and Remediation Guide

SA
Security Arsenal Team
September 29, 2026
13 min read

NVD has published CVE-2026-103040, a critical vulnerability rated CVSS 9.8 (CRITICAL, network-exploitable) in LightLLM, the Python-based LLM inference and serving framework that has become a common building block in AI/ML production infrastructure. LightLLM versions through 1.2.0 contain an unauthenticated remote code execution vulnerability in the router profiler service that is activated when the router is started with the --enable_profiling flag.

The flaw is straightforward and severe: when profiling is enabled, the service exposes an unauthenticated RPyC (Remote Python Call) server with pickle deserialization enabled. Any attacker who can reach the profiler's TCP port can submit crafted serialized Python objects to the profiler command queue and achieve arbitrary code execution on the host running the inference router — with the privileges of the LightLLM process, which in many deployments runs inside GPU-backed containers with broad access to model weights, training data, API keys, and internal service meshes.

A CVSS 9.8 score is appropriate here: the vulnerability requires no authentication, no user interaction, and low attack complexity, and it results in complete compromise of confidentiality, integrity, and availability. Insecure pickle deserialization is one of the most reliably exploitable bug classes in the Python ecosystem — public deserialization exploitation primitives are mature, weaponization is trivial, and AI serving infrastructure is increasingly a first-class target for both opportunistic attackers and nation-state operators seeking access to proprietary models and data.

If your organization runs LightLLM in production, staging, or even development environments reachable beyond localhost, treat this as an immediate patching and exposure audit event.

Technical Analysis

Affected Products and Versions

ItemDetail
ProductLightLLM (open-source LLM inference/serving framework)
Affected versionsAll versions through 1.2.0
Vulnerable componentRouter profiler service
Trigger conditionRouter started with the --enable_profiling flag
CVECVE-2026-103040
CVSS9.8 (CRITICAL) — Vector: network, no auth, no user interaction
Vulnerability classCWE-502: Deserialization of Untrusted Data

The critical scoping fact: the vulnerable profiler service only exists when --enable_profiling is passed at router startup. LightLLM instances started without that flag are not exposed to this specific attack surface — but profiling flags are routinely enabled during performance tuning, capacity planning, and debugging, and they frequently persist in deployment manifests, Helm values, systemd units, and container entrypoints long after the original tuning exercise ended.

How the Vulnerability Works

  1. Startup with profiling enabled. The LightLLM router process starts an auxiliary profiler service implemented on RPyC, a Python library for transparent remote procedure calls. RPyC's classic (server) mode executes client-supplied Python objects on the server side.
  2. No authentication layer. The RPyC server is exposed without authentication — any TCP client that can reach the port can connect and interact with it. There is no credential check, no allowlist, and no mTLS in the default configuration.
  3. Pickle deserialization enabled. The profiler command queue accepts pickle-serialized objects. Python's pickle module is explicitly documented as unsafe for untrusted input: during deserialization, pickle can invoke arbitrary callables embedded in the serialized stream (the classic __reduce__ pattern), which gives the attacker direct code execution the moment the object is unpickled.
  4. Code execution as the LightLLM process. The attacker's payload executes in the context of the LightLLM router process. In typical deployments this means: access to loaded model weights, tokenizer files, environment variables containing API keys (OpenAI, Hugging Face, cloud provider credentials), Kubernetes service account tokens, and — critically — GPU-equipped hosts that are often on flat internal networks with little egress filtering.

Why This Is Worse Than a Typical RCE

AI inference infrastructure sits at an unusual intersection of value and exposure:

  • Model weights are crown-jewel IP. A successful intrusion enables direct exfiltration of fine-tuned models.
  • GPU hosts are expensive and often under-monitored. Many organizations deploy LLM serving stacks outside their traditional EDR coverage, sometimes as bare Python processes inside containers with no runtime security agent.
  • Profiling flags create hidden exposure. Operators enable --enable_profiling during benchmarking and forget to disable it. The vulnerable service may be listening on production routers right now without anyone on the security team knowing it exists.
  • RPyC is a known-risky protocol. Any RPyC classic server exposed to a network should be treated as pre-authenticated code execution by design, whether or not this CVE exists.

Exploitation Status

As of publication, CVE-2026-103040 has been published by NVD with a CVSS 9.8 rating. The vulnerability description and affected component are public, which means exploitation details are derivable by any competent attacker — pickle deserialization exploitation is a solved problem with abundant public tooling. There is no confirmed CISA KEV listing at the time of writing, but defenders should assume that weaponized exploitation is imminent or already occurring, particularly given the rapid targeting of AI/ML infrastructure observed across 2025–2026 threat reporting. Treat internet-exposed LightLLM profiler services as compromised until proven otherwise.

Detection & Response

The following detection content targets the specific, observable behaviors of this vulnerability: the presence of the --enable_profiling flag in LightLLM process command lines, unauthenticated inbound connections to RPyC-style services from LightLLM hosts, and post-exploitation behavior such as the LightLLM router process spawning unexpected child processes.

Sigma Rules

YAML
---
title: LightLLM Router Started with Profiling Flag (CVE-2026-103040 Exposure)
id: a1b2c3d4-1111-4e2a-9b3c-0d1e2f3a4b5c
status: experimental
description: Detects LightLLM server or router processes started with --enable_profiling, which exposes the unauthenticated RPyC profiler service vulnerable to CVE-2026-103040.
references:
  - https://nvd.nist.gov/vuln/detail/CVE-2026-103040
author: Security Arsenal
date: 2026/04/06
tags:
  - attack.execution
  - attack.t1190
logsource:
  category: process_creation
  product: linux
detection:
  selection_proc:
    CommandLine|contains:
      - 'lightllm'
  selection_flag:
    CommandLine|contains:
      - '--enable_profiling'
  condition: selection_proc and selection_flag
falsepositives:
  - Intentional profiling during controlled benchmarking in isolated environments
level: high
---
title: LightLLM Process Spawning Unexpected Child Process (Possible CVE-2026-103040 Exploitation)
id: a1b2c3d4-2222-4e2a-9b3c-0d1e2f3a4b5c
status: experimental
description: Detects LightLLM router/server processes spawning shell or scripting interpreters, consistent with post-exploitation activity following unauthenticated RPyC/pickle deserialization RCE.
references:
  - https://nvd.nist.gov/vuln/detail/CVE-2026-103040
author: Security Arsenal
date: 2026/04/06
tags:
  - attack.execution
  - attack.t1059
logsource:
  category: process_creation
  product: linux
detection:
  selection_parent:
    ParentCommandLine|contains:
      - 'lightllm'
  selection_child:
    CommandLine|contains:
      - '/bin/sh'
      - '/bin/bash'
      - '/bin/dash'
      - 'python -c'
      - 'curl '
      - 'wget '
      - 'ncat'
      - 'nc -'
      - '/tmp/'
  condition: selection_parent and selection_child
falsepositives:
  - Legitimate operational scripts invoked by administrators during profiling sessions
level: critical

Analyst guidance on these rules: the first rule is an exposure rule — it will (correctly) fire on every LightLLM instance running with profiling enabled. That is the point: it builds your inventory of vulnerable systems. The second rule is a behavioral post-exploitation rule — LightLLM routers have essentially no legitimate reason to spawn shells or download utilities, so fidelity is high. If you see lightllm spawning bash or curl to an external host, treat it as an active intrusion, not a tuning exercise.

KQL — Microsoft Sentinel / Defender

LightLLM runs on Linux; the queries below cover both Defender for Endpoint telemetry (DeviceProcessEvents/DeviceNetworkEvents) and Syslog/CEF ingestion for hosts without MDE.

KQL — Microsoft Sentinel / Defender
// Hunt 1: LightLLM processes running with --enable_profiling (exposure inventory)
union isfuzzy=true
( DeviceProcessEvents
  | where InitiatingProcessCommandLine has "lightllm" or ProcessCommandLine has "lightllm"
  | where ProcessCommandLine has "--enable_profiling"
  | project TimeGenerated, DeviceName, ProcessCommandLine, AccountName, InitiatingProcessFileName, Source = "MDE" ),
( Syslog
  | where ProcessName has "lightllm" or SyslogMessage has_all ("lightllm", "--enable_profiling")
  | project TimeGenerated, Computer, ProcessName, SyslogMessage, Source = "Syslog" )
| sort by TimeGenerated desc;

// Hunt 2: Post-exploitation - lightllm spawning shells, interpreters, or downloaders
DeviceProcessEvents
| where InitiatingProcessCommandLine has "lightllm"
| where FileName in~ ("sh", "bash", "dash", "python", "python3", "perl", "curl", "wget", "nc", "ncat", "socat")
   or ProcessCommandLine has_any ("/bin/sh", "/bin/bash", "-c ", "base64", "/dev/tcp")
| project TimeGenerated, DeviceName, InitiatingProcessCommandLine, FileName, ProcessCommandLine, AccountName, RemoteIP
| sort by TimeGenerated desc;

// Hunt 3: Inbound network connections to suspected RPyC/profiler listeners from external or unusual sources
DeviceNetworkEvents
| where InitiatingProcessCommandLine has "lightllm" or LocalPort in (18861, 18812)
| where ActionType == "InboundConnectionAccepted" or RemoteIPType == "Public"
| extend IsExternal = iff(ipv4_is_private(RemoteIP), "Internal", "External")
| project TimeGenerated, DeviceName, LocalPort, RemoteIP, IsExternal, RemoteUrl, InitiatingProcessCommandLine
| sort by TimeGenerated desc;

Adjust the port list in Hunt 3 to match your environment — the profiler listener port is configuration-dependent, so the lightllm process-name match is the more reliable pivot. Any public inbound connection accepted by a LightLLM process outside your documented API port warrants immediate investigation.

Velociraptor VQL

VQL — Velociraptor
-- Hunt: LightLLM routers with profiling enabled and their listening sockets (CVE-2026-103040)
-- Combines process command-line inspection with live network state per host.
LET procs = SELECT Pid, Name, CommandLine, Exe, Username, CreateTime
FROM pslist()
WHERE CommandLine =~ '(?i)lightllm'

LET flagged = SELECT Pid, Name, CommandLine, Username, CreateTime,
       if(condition=CommandLine =~ '(?i)--enable_profiling',
          then='VULNERABLE_FLAG_PRESENT', else='profiling_not_enabled') AS ProfilingStatus
FROM procs

LET conns = SELECT Pid, Fd, Family, Type, Laddr, Lport, Raddr, Rport, Status, Name
FROM netstat()
WHERE Status =~ '(?i)listen' AND Name =~ '(?i)python|lightllm'

SELECT f.Pid, f.Name, f.Username, f.CreateTime, f.ProfilingStatus, f.CommandLine,
       c.Laddr, c.Lport, c.Status
FROM flagged AS f
LEFT JOIN conns AS c ON f.Pid = c.Pid

Deploy this as a hunt across your GPU/compute fleet. Any row where ProfilingStatus = 'VULNERABLE_FLAG_PRESENT' and the process has a listener bound to 0.0.0.0 or a routable interface is an immediately exploitable system. If the hunt returns nothing, verify the fleet actually runs LightLLM before concluding you are unaffected.

Remediation / Hardening Script

The following Bash script audits Linux hosts for running LightLLM processes with the profiling flag, checks common deployment artifacts (systemd units, Docker containers, Kubernetes-style manifests), and — with the --kill option — terminates exposed routers so they can be restarted without the flag.

Bash / Shell
#!/usr/bin/env bash
# CVE-2026-103040 exposure audit: LightLLM --enable_profiling
# Usage: sudo ./cve-2026-103040-audit.sh [--kill]
set -euo pipefail

KILL_MODE=0
[[ "${1:-}" == "--kill" ]] && KILL_MODE=1

echo "=== CVE-2026-103040: LightLLM profiler exposure audit ==="
echo

# 1. Running LightLLM processes with --enable_profiling
echo "[1] Running LightLLM processes:"
EXPOSED_PIDS=$(pgrep -af 'lightllm' | grep -- '--enable_profiling' | awk '{print $1}' || true)
if [[ -n "$EXPOSED_PIDS" ]]; then
  pgrep -af 'lightllm' | grep -- '--enable_profiling'
  echo "  [!] EXPOSED: profiling flag active on PID(s): $EXPOSED_PIDS"
else
  echo "  [OK] No running lightllm process with --enable_profiling found."
fi
echo

# 2. Listening sockets owned by lightllm/python processes
echo "[2] Listening sockets for lightllm/python processes:"
ss -tlnp 2>/dev/null | grep -Ei 'lightllm|python' || echo "  [OK] No lightllm/python listeners detected."
echo

# 3. Persistent configuration references to the flag
echo "[3] Searching systemd units, compose files, and manifests for --enable_profiling:"
grep -RIn -- '--enable_profiling' \
  /etc/systemd/system /usr/lib/systemd/system \
  /opt /srv /home/*/docker-compose.y*ml /etc/docker 2>/dev/null \
  || echo "  [OK] No persistent config references found in scanned paths."
echo

# 4. Egress check: any established non-API connections from lightllm processes
echo "[4] Established connections from lightllm processes (review for C2/exfil):"
ss -tnp 2>/dev/null | grep -Ei 'lightllm|python' | grep -v ':6379\|:5432\|:8000' || echo "  [OK] None."
echo

# 5. Optional: terminate exposed routers (they must be restarted WITHOUT the flag)
if [[ $KILL_MODE -eq 1 && -n "$EXPOSED_PIDS" ]]; then
  echo "[5] --kill specified: terminating exposed router PID(s): $EXPOSED_PIDS"
  kill -TERM $EXPOSED_PIDS
  sleep 3
  pgrep -af 'lightllm' | grep -- '--enable_profiling' && kill -KILL $EXPOSED_PIDS || true
  echo "  [DONE] Restart the router from your deployment tooling WITHOUT --enable_profiling."
else
  echo "[5] Remediation: remove --enable_profiling from startup args and restart, or run with --kill to terminate now."
fi
echo
echo "=== Audit complete. Upgrade LightLLM to a fixed release as soon as one is available. ==="

Run this on every GPU host, inference node, and container host in your fleet. In Kubernetes environments, additionally grep your manifests, Helm values, and ConfigMaps for enable_profiling and enable-profiling, and check container entrypoints with kubectl get pods -o jsonpath for the flag.

Remediation

CVE-2026-103040 is a configuration-triggered vulnerability, which gives defenders two independent and complementary lines of defense. Execute both.

1. Immediate: Disable the Profiling Flag (Primary Workaround)

  • Remove --enable_profiling from every LightLLM router startup invocation — systemd units, Docker/Compose entrypoints, Kubernetes manifests, Helm values, supervisor configs, and CI/CD deployment templates.
  • Restart affected routers. The vulnerable RPyC service only exists while the process is running with the flag; a restart without it eliminates the attack surface.
  • Audit for shadow enablement. The flag frequently survives in staging manifests that get promoted to production, or in commented-but-active config files. The detection script and the Sigma exposure rule above are designed to find exactly this.

2. Immediate: Network Containment

  • Block inbound access to the profiler listener. Identify the port the profiler binds to on your hosts (see the VQL artifact and ss -tlnp output) and restrict it to localhost or a dedicated management network via host firewall (nftables/iptables) and security groups.
  • Deny public ingress to LightLLM hosts on any port other than your documented inference API endpoint. AI serving nodes should never accept arbitrary inbound TCP from the internet.
  • Restrict egress from GPU hosts. Post-exploitation of this class of vulnerability depends on outbound connectivity for payload staging and exfiltration. Egress filtering to an allowlist materially degrades attacker follow-on activity.

3. Upgrade Path

  • Monitor the official LightLLM project (GitHub: ModelTC/lightllm) and the NVD entry — https://nvd.nist.gov/vuln/detail/CVE-2026-103040 — for the fixed release superseding 1.2.0, and upgrade as soon as it is published. At time of writing, the advisory describes all versions through 1.2.0 as affected, so flag removal + network containment is the only complete mitigation until a patch lands.
  • Subscribe to the project's security advisories and add LightLLM to your SBOM-driven vulnerability monitoring (e.g., Grype/Trivy scans of container images) so the fix version triggers automatically in your pipeline.

4. Threat-Hunt Before and After Remediation

Because exploitation requires no authentication and leaves minimal host artifacts, do not assume a clean state just because you disabled the flag:

  • Run the KQL Hunts 2 and 3 and the Velociraptor artifact across all LightLLM hosts, covering at least the last 30 days (or maximum telemetry retention).
  • Look specifically for: LightLLM processes spawning shells or Python one-liners, unexpected outbound connections from inference hosts, new files in /tmp and /dev/shm, new systemd units or cron entries, and unexpected accounts or SSH keys.
  • If any post-exploitation indicator is found, treat it as a full incident: isolate the host, capture memory if feasible, rotate all credentials accessible to the LightLLM environment (API keys, cloud credentials, Kubernetes service account tokens), and assess whether model weights or training data were exfiltrated.

5. Strategic Hardening for AI/ML Infrastructure

This CVE is a symptom of a broader pattern: ML serving stacks ship with powerful debug and profiling surfaces that are unsafe on any network. Fold these controls into your platform baseline:

  • Policy-as-code guardrails that reject any deployment manifest containing profiling/debug flags for ML serving frameworks.
  • Runtime security coverage (Falco, eBPF-based sensors, or MDE for Linux) on GPU hosts — these machines are too often outside EDR scope.
  • Network segmentation placing inference infrastructure in its own zone with default-deny inbound and tightly allowlisted egress.
  • Inventory discipline: you cannot defend LightLLM instances you do not know exist. Include ML frameworks in attack surface management and external scanning.

Related Resources

Security Arsenal Penetration Testing Services AlertMonitor Platform Book a SOC Assessment vulnerability-management Intel Hub

Is your security operations ready?

Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.

CVE-2026-103040: LightLLM Profiler Service Unauthenticated RCE (CVSS 9.8) — Detection and Remediation Guide | Security Arsenal | Security Arsenal