NVD has published CVE-2026-103041, a CVSS 9.8 CRITICAL vulnerability affecting LightLLM through version 1.2.0 in multimodal deployments. The flaw is as dangerous as it is simple: LightLLM exposes an unauthenticated RPyC (Remote Python Call) cache service — with Python pickle deserialization enabled — bound to all network interfaces. Any attacker who can reach that port over the network can submit a crafted serialized object and execute arbitrary code with the privileges of the LightLLM service. No credentials. No user interaction. No exploit chain required beyond a malicious pickle payload.
If your organization runs self-hosted LLM inference — and in 2026, an enormous number of enterprises do, often outside traditional change control — this is a drop-everything remediation item. AI inference servers sit on some of the most valuable real estate in your environment: GPU-rich hosts with access to model weights, training data, vector stores, and frequently broad internal network reach. A trivially exploitable, unauthenticated RCE on these systems is a direct path to data theft, model exfiltration, cryptomining on expensive GPU capacity, and lateral movement.
Technical Analysis
Affected Products and Versions
- Product: LightLLM (lightllm), a Python-based LLM inference and serving framework
- Affected versions: All versions through 1.2.0
- Affected configurations: Multimodal deployments, which instantiate the RPyC-based cache service used to share multimodal data (images, embeddings, intermediate tensors) between components
- Attack vector: Network (CVSS:3.1 AV:N — remote, unauthenticated)
- CVE / Score: CVE-2026-103041 — CVSS 9.8 CRITICAL
- Advisory: https://nvd.nist.gov/vuln/detail/CVE-2026-103041
How the Vulnerability Works
RPyC is a transparent Python RPC library. Its classic service model is inherently trusting: by design, it allows remote clients to execute Python operations as if they were local. When a service additionally exposes pickle-based deserialization of attacker-supplied objects, the attacker controls the object graph being reconstructed inside the server process.
Pickle deserialization in Python is well understood to be equivalent to code execution when the input is untrusted. A crafted pickle stream can invoke arbitrary callables during __reduce__ reconstruction — the canonical example being os.system or subprocess.Popen invoked with attacker-controlled arguments. In this case:
- Exposure: The LightLLM multimodal cache service binds to
0.0.0.0(all interfaces), not127.0.0.1. Any host with network reachability — including from the internet if firewall policy permits, or from any compromised internal host — can connect. - No authentication: The RPyC service accepts connections without credentials, tokens, or mutual TLS.
- Deserialization: Exposed cache methods accept serialized objects. Pickle deserialization is enabled.
- Execution: The attacker sends a crafted serialized object; the payload executes in the context of the LightLLM service account, which typically has access to GPU devices, model files, environment variables containing API keys and cloud credentials, and whatever network segments the host can reach.
This is the same architectural anti-pattern that has burned the Python ML ecosystem repeatedly: pickle treated as a data format rather than as what it actually is — an executable serialization protocol.
Exploitation Status
At the time of publication, CVE-2026-103041 is published in NVD with a network-exploitable pathway and no authentication requirement — the lowest possible exploitation bar. Given that exploitation requires only network reachability and a well-known payload technique (malicious pickle objects are a solved problem with public tooling), defenders should treat exploitation as imminent and prioritize as if active. RPyC exposure is also trivially discoverable via internet scanning (Shodan/Censys), so any internet-facing instance should be presumed enumerated within days of disclosure. Monitor CISA KEV for inclusion.
Detection & Response
The highest-fidelity detection signals for this threat are: (1) the LightLLM Python process spawning unexpected child processes (the pickle payload executing a shell or downloader), (2) inbound network connections to the RPyC cache service port from non-LightLLM hosts, and (3) the cache service listening on a non-loopback interface at all.
---
title: LightLLM or Python Inference Service Spawning Shell or System Utility
id: 9c4e1a72-5b3d-4f8a-bc61-2e7d9a0f4312
status: experimental
description: Detects shell interpreters, downloaders, or reconnaissance utilities spawned as child processes of LightLLM or generic Python inference server processes — a strong indicator of successful pickle deserialization RCE (CVE-2026-103041).
references:
- https://nvd.nist.gov/vuln/detail/CVE-2026-103041
- https://attack.mitre.org/techniques/T1059/
author: Security Arsenal
date: 2026/04/06
tags:
- attack.execution
- attack.t1059.004
- attack.t1059.006
logsource:
category: process_creation
product: linux
detection:
selection_parent:
ParentImage|endswith:
- '/python'
- '/python3'
- '/python3.10'
- '/python3.11'
- '/python3.12'
ParentCommandLine|contains:
- 'lightllm'
- 'rpyc'
- 'lightllm.server'
selection_child:
Image|endswith:
- '/sh'
- '/bash'
- '/dash'
- '/zsh'
- '/curl'
- '/wget'
- '/nc'
- '/ncat'
- '/python'
- '/python3'
- '/base64'
- '/chmod'
- '/id'
- '/whoami'
- '/uname'
condition: selection_parent and selection_child
falsepositives:
- LightLLM health-check wrappers or operator scripts invoked by the serving process in custom deployments
level: critical
---
title: RPyC Cache Service Bound to Non-Loopback Interface
id: 3f8b2d61-7a04-4c19-9e53-8d1c6b5a9204
status: experimental
description: Detects Python/LightLLM processes establishing network listeners on 0.0.0.0 or external interfaces consistent with the exposed unauthenticated RPyC cache service in CVE-2026-103041. Review any LightLLM listener not bound to 127.0.0.1.
references:
- https://nvd.nist.gov/vuln/detail/CVE-2026-103041
- https://attack.mitre.org/techniques/T1190/
author: Security Arsenal
date: 2026/04/06
tags:
- attack.initial_access
- attack.t1190
logsource:
category: network_connection
product: linux
detection:
selection:
Image|endswith:
- '/python'
- '/python3'
CommandLine|contains:
- 'lightllm'
- 'rpyc_classic'
- 'rpyc'
DestinationIp:
- '0.0.0.0'
- '::'
Initiated: 'false'
condition: selection
falsepositives:
- Legitimate LightLLM API server listeners (validate port ownership; the cache service should never be externally bound)
level: high
---
title: Inbound Connection to RPyC Service Port from Non-LLM Host
id: b71e6c48-2d95-4f30-a814-6c3a9f2e7851
status: experimental
description: Detects network connections to the RPyC cache service port on LightLLM hosts originating from hosts outside the approved inference cluster — potential exploitation of CVE-2026-103041.
references:
- https://nvd.nist.gov/vuln/detail/CVE-2026-103041
- https://attack.mitre.org/techniques/T1190/
author: Security Arsenal
date: 2026/04/06
tags:
- attack.initial_access
- attack.t1190
logsource:
category: firewall
product: linux
detection:
selection:
DestinationPort:
- 18812
- 18861
- 18878
Action: 'allowed'
filter_loopback:
SourceIp|startswith:
- '127.'
- '::1'
condition: selection and not filter_loopback
falsepositives:
- Intra-cluster LightLLM component communication — baseline approved cluster CIDRs and alert on all other sources
level: high
Note on ports: RPyC classic servers commonly use TCP 18812; LightLLM deployments may configure different ports for the cache service. Inventory your actual deployment (ss -tlnp on inference hosts) and tune the port list to your environment — a rule on the wrong port is silent, not safe.
KQL — Microsoft Sentinel / Defender
This query hunts for the exploitation pattern across Linux inference hosts ingested via Syslog/CEF, plus Defender for Endpoint process telemetry where deployed on GPU servers:
// Hunt 1: Python/LightLLM inference processes spawning shells, downloaders, or recon tools
let SuspiciousChildren = dynamic(["/bin/sh", "/bin/bash", "/usr/bin/curl", "/usr/bin/wget", "/bin/nc", "/usr/bin/ncat", "/usr/bin/base64", "/usr/bin/id", "/usr/bin/whoami", "/usr/bin/uname", "/bin/chmod"]);
union isfuzzy=true
(DeviceProcessEvents
| where InitiatingProcessFileName in~ ("python", "python3", "python3.10", "python3.11", "python3.12")
| where InitiatingProcessCommandLine has_any ("lightllm", "rpyc")
| where FileName in~ ("sh", "bash", "curl", "wget", "nc", "ncat", "base64", "id", "whoami", "uname", "chmod")
| project TimeGenerated=TimeGenerated, DeviceName, InitiatingProcessCommandLine, FileName, ProcessCommandLine, AccountName, Source="MDE"),
(Syslog
| where ProcessName has_any ("python", "lightllm")
| where SyslogMessage has_any ("/bin/sh", "/bin/bash", "curl ", "wget ", "nc -", "base64 -d")
| project TimeGenerated, DeviceName=HostName, InitiatingProcessCommandLine=ProcessName, FileName="", ProcessCommandLine=SyslogMessage, AccountName="", Source="Syslog")
| order by TimeGenerated desc;
// Hunt 2: Inbound connections to RPyC cache ports from non-loopback sources (CEF/Syslog firewall data)
CommonSecurityLog
| where DestinationPort in (18812, 18861, 18878)
| where not(SourceIP startswith "127.")
| summarize ConnectionCount=count(), FirstSeen=min(TimeGenerated), LastSeen=max(TimeGenerated), SourceIPs=make_set(SourceIP, 20) by DestinationIP, DestinationPort, DeviceAction
| order by ConnectionCount desc;
Velociraptor VQL
Use this artifact to sweep inference hosts for the exposed listener and suspicious process lineage in one pass:
-- Hunt for exposed RPyC/LightLLM listeners on non-loopback interfaces
-- and suspicious child processes of Python inference services
SELECT Pid, Name, CommandLine, Address, Port, Status
FROM netstat()
WHERE (Name =~ 'python' OR CommandLine =~ '(?i)lightllm|rpyc')
AND Status = 'LISTEN'
AND NOT Address =~ '^(127\.|::1|\[::1\])'
UNION ALL
SELECT Pid, Name, CommandLine, '' AS Address, 0 AS Port, 'CHILD_PROC' AS Status
FROM pslist()
WHERE CommandLine =~ '(?i)(/bin/(ba)?sh|curl |wget |nc -|base64 -d)'
AND Ppid IN (
SELECT Pid FROM pslist()
WHERE CommandLine =~ '(?i)lightllm|rpyc'
)
Verification & Hardening Script
Run this Bash script across your inference fleet to identify exposed RPyC listeners, check LightLLM version, and apply compensating controls pending patching:
#!/bin/bash
# CVE-2026-103041 - LightLLM RPyC exposure verification and interim hardening
set -uo pipefail
echo "=== [1/4] LightLLM version check ==="
if command -v pip3 >/dev/null 2>&1; then
pip3 show lightllm 2>/dev/null | grep -E '^(Name|Version)' || echo "lightllm not found via pip3"
python3 -c "import lightllm; print('lightllm', lightllm.__version__)" 2>/dev/null || true
fi
echo "=== [2/4] Checking for Python listeners on non-loopback interfaces ==="
ss -tlnp 2>/dev/null | grep -E 'python|lightllm' | grep -vE '127\.0\.0\.1|\[::1\]' || echo "No externally bound python listeners found"
echo "=== [3/4] Checking for rpyc processes and common RPyC ports ==="
ps aux | grep -iE 'rpyc|lightllm' | grep -v grep || echo "No lightllm/rpyc processes found"
ss -tlnp 2>/dev/null | awk '$4 ~ /:(18812|18861|18878)$/ {print "EXPOSED PORT:", $0}'
echo "=== [4/4] Applying interim iptables hardening (block external RPyC access) ==="
read -rp "Apply host firewall rules to block external access to RPyC cache ports? [y/N] " CONFIRM
if [[ "$CONFIRM" =~ ^[Yy]$ ]]; then
for PORT in 18812 18861 18878; do
# Allow loopback, drop everything else to cache ports
iptables -C INPUT -i lo -p tcp --dport "$PORT" -j ACCEPT 2>/dev/null || \
iptables -A INPUT -i lo -p tcp --dport "$PORT" -j ACCEPT
iptables -C INPUT -p tcp --dport "$PORT" -j DROP 2>/dev/null || \
iptables -A INPUT -p tcp --dport "$PORT" -j DROP
echo "Hardened port $PORT (loopback only)"
done
echo "NOTE: Persist rules with iptables-save / netfilter-persistent per your distro."
else
echo "Skipped firewall changes."
fi
echo "=== Done. Review output above. Any externally bound cache listener = CVE-2026-103041 exposure. ==="
Remediation
-
Upgrade LightLLM immediately. Update to a release later than 1.2.0 that remediates the exposure — verify the patched version against the upstream LightLLM repository and the NVD advisory (https://nvd.nist.gov/vuln/detail/CVE-2026-103041). Confirm via
pip3 show lightllmafter upgrade and re-test that the cache service no longer binds externally or accepts unauthenticated pickle payloads. -
Bind the cache service to loopback or disable it. If patching cannot happen today, configure the cache service to listen on
127.0.0.1only, or disable the multimodal cache service if your workload does not require it. Treat any RPyC service bound to0.0.0.0as a critical misconfiguration independent of this CVE. -
Network segmentation as a hard control. Place inference hosts in a dedicated VLAN/security group with ingress restricted to the API gateway tier only. No host other than approved LightLLM cluster members should have any path to RPyC ports. Block internet egress from inference hosts except to an explicit allowlist (model registries, package mirrors) to blunt post-exploitation C2 and exfiltration.
-
Hunt before you patch. Because this vulnerability is trivially exploitable and scanners will find exposed instances quickly, run the KQL and VQL hunts above across a retro window of at least 30 days before assuming a clean state. Any shell spawned by the inference process, any unexpected connection to a cache port, or any new files/cron entries/systemd units on inference hosts warrants full IR triage.
-
Rotate credentials on inference hosts. If you confirm an exposed listener existed prior to remediation, assume compromise: rotate cloud IAM keys, model registry tokens, API keys in environment variables, and any secrets reachable from the service account.
-
Fix the architectural pattern. Audit every Python service in your environment for pickle deserialization of network-supplied data and unauthenticated RPyC endpoints. Replace pickle with safe serialization (JSON, MessagePack with strict schemas) for any data crossing a trust boundary. Add RPyC listener detection to your continuous configuration monitoring — this will not be the last ML-framework RCE of this shape.
Related Resources
Security Arsenal Penetration Testing Services AlertMonitor Platform Book a SOC Assessment vulnerability-management Intel Hub
Is your security operations ready?
Get a free SOC assessment or see how AlertMonitor cuts through alert noise with automated triage.