LMCache CVE-2026-105192: unauthenticated RCE in your LLM cache, no patch yet
LMCache CVE-2026-105192 is an unauthenticated RCE in the LLM KV-cache behind vLLM: one packet to the ZeroMQ port runs code as root. No patch yet.
Published by Zero Hunt, an autonomous AI red team on an on-premise appliance running private AI: automated penetration testing for networks and infrastructure, black-box or gray-box, with a human approving every step that matters.
If you run self-hosted LLM inference at any scale, you almost certainly run a KV-cache layer in front of it, and there is a good chance it is LMCache — the open-source cache that sits between vLLM workers and holds the key-value tensors that make long-context and multi-turn serving affordable. On October 7, 2026 JFrog's security research team disclosed CVE-2026-105192, an unauthenticated remote code execution flaw in LMCache's multiprocess transport. A single crafted network message to the cache's ZeroMQ port runs attacker code with the privileges of the process — root, in the official containers. There is no fixed release. The only thing standing between that port and an attacker is the network you put it on.
Developing story — first published 19:28 CEST (17:28 UTC), October 7, 2026. Updated as JFrog, the LMCache maintainers and CISA publish more.
At a glance
| CVE | CVE-2026-105192 |
| Product / affected versions | LMCache multiprocess mode · 0.3.9 (October 2025) through 0.5.5 · 0.5.6 release candidates · development branch |
| Fixed in | No fixed release as of 2026-10-07 |
| CVSS | 9.8 Critical · CVSS 3.1 AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H · scored by JFrog; NVD published 2026-10-07 (CWE-306; CWE-502) |
| Exploited in the wild | No evidence as of 2026-10-07 (JFrog) |
| CISA KEV | Not listed (catalog version 2026.10.04) |
| Public PoC | Not released with the advisory (JFrog, 2026-10-07) |
| Official advisory | JFrog JFSA-2026-001694382 |
What CVE-2026-105192 actually is
LMCache can run its cache as a standalone server that vLLM workers reach over ZeroMQ — the "multiprocess mode" that lets you scale cache capacity independently of the model workers. The transport is an unauthenticated ZeroMQ ROUTER socket on port 5555. Workers send typed messages; the server decodes them with msgpack.
The bug is in how one message type is decoded. The REGISTER_KV_CACHE message carries a msgpack extension (code 1) whose payload is handed to pickle.loads inside DeviceIPCWrapper.Deserialize. Pickle is a Python serialization format that can carry executable instructions and run them as the stream is decoded — it is the canonical unsafe-deserialization sink (CWE-502). The server unpacks the pickled object while it is still reading the message's arguments, before any check of the message type. So a crafted message does not need to be a valid registration; it only needs to reach the socket. When it is decoded, the sender's code runs.
There is no authentication on the socket (CWE-306), so "reach the socket" is the whole exploit. JFrog scored it CVSS 9.8 for the network-exposed configuration.
The one piece of good news is the default: LMCache binds the multiprocess transport to localhost unless an operator sets a routable address. The problem is that multi-node deployments — the ones with the biggest caches and the most GPUs behind them — do exactly that.
Who is exposed
You are in scope if all of these are true:
- You run LMCache 0.3.9 or later (any release up to and including 0.5.5, the 0.5.6 release candidates, or a build off the development branch).
- You run it in multiprocess mode (a standalone cache server, not the in-process library).
- The multiprocess port (default 5555) is reachable from anywhere you do not fully trust — because an operator passed
--host 0.0.0.0or another routable address, or because a container/orchestrator published the port, or because "trusted cluster network" turned out to include a pod subnet an attacker can reach.
Check it directly. On the cache host, find what the transport is bound to and who can reach it:
# Is the LMCache multiprocess port listening on a routable address?
ss -ltnp | grep -E ':5555|lmcache'
# Grep your launch config / orchestrator manifests for a routable host
grep -RniE 'lmcache|--host|LMCACHE_.*HOST|5555' /etc/ deploy/ k8s/ 2>/dev/null
# From another host on the same network segment, can you open the port?
nc -vz <cache-host> 5555
A 0.0.0.0:5555 in the first command, or a successful connect from the third, means the cache is one packet away from code execution and there is no patch to install.
Remediation
There is no fixed version, so remediation here is not "patch and move on." It is a set of compensating controls you apply now and re-verify, in priority order.
1. Am I affected? Run the three checks above. If the port is on localhost only and multiprocess mode is not exposed, you are not reachable over the network — hold at the monitoring step and wait for a fixed release.
2. Get the port off the network — this is the fix until there is a patch. Per JFrog: do not set --host to a routable address; keep the multiprocess port on localhost or on a trusted, isolated cluster network. If you do not need multiprocess mode, disable it and run the in-process cache.
3. Can't change the topology right now? — firewall the socket. Restrict port 5555 to the exact worker IPs that must talk to the cache, at the host firewall and the network layer both. In Kubernetes, a default-deny NetworkPolicy with an explicit allow from the vLLM worker pods only. Treat "internal network" as untrusted: an unauthenticated pickle sink does not distinguish a worker from anything else that can open the socket.
4. Hunt for compromise. JFrog reports no in-the-wild exploitation as of 2026-10-07, and no vendor IOCs have been published — so hunt on behavior, not signatures. On any cache host that has been network-exposed, look for child processes spawned by the LMCache/Python process (shells, curl/wget, interpreters), outbound connections from a host that should only serve cache traffic, and new ZeroMQ peers on 5555 that are not your known workers. Map it to MITRE ATT&CK T1190 (exploit public-facing application) and T1059 (command/scripting interpreter). Because code runs as root in the official containers, a successful exploit owns the container and whatever the container can reach — including the model weights and any credentials mounted into it.
5. Eradicate and verify. A compromised cache node should be rebuilt, not cleaned — pickle RCE gives an attacker arbitrary first-stage code and you cannot enumerate what it did. Rotate any secret that lived on or was reachable from the node: model-registry tokens, object-store keys, service-account tokens. Then confirm the port is no longer reachable from untrusted networks, and keep watching 5555 until you are running a fixed release.
What is not known yet
- No fixed release. As of this writing there is no patched LMCache version; the maintainers' durable fix will be to stop passing network bytes to pickle (replace the serializer behind msgpack extension code 1, add CURVE or HMAC authentication on the transport). Track the LMCache repository for the release.
- No CISA KEV listing (catalog 2026.10.04) and no confirmed exploitation. That is the window, not the all-clear: the mechanism is public, the sink is a one-line pickle call, and a working exploit is not hard to build from the advisory.
- Real-world exposure count is unknown. How many LMCache deployments run multiprocess mode on a routable address is not published. If you operate one, you know the answer for your own fleet after the checks above — most teams have never looked.
The asset nobody inventoried
The honest problem CVE-2026-105192 exposes is not pickle. It is that the LLM-serving stack grew faster than anyone's asset register. A cache server on port 5555, stood up to make inference cheaper, is not in the CMDB, not in the scanner's scope, and not on the list of things a quarterly pentest looks at. A signature scanner sees a ZeroMQ ROUTER and a listening port; it does not know that one message type deserializes with pickle before checking its own type. The gap is not the bug — it is that no one was testing the inference tier as an attack surface at all.
That gap is exactly what an autonomous AI red team closes, and it is the operational question this disclosure raises: with no patch available, what do I do now, and in what order? Zero Hunt's AI Remediation Advisor is built for that moment — it ranks findings by real exploitability (CISA KEV first, then what was proven reachable on your own network, then CVSS and EPSS), and for an unpatchable flaw it gives the compensating control that actually removes the exposure — here, getting the ZeroMQ socket off any routable interface — with the rollout and rollback notes to apply it safely, then re-verifies by re-running the same proof that demonstrated the issue. A version string is a claim; a re-run is evidence.
Upstream of that, the reason the Advisor has something to rank is that the 10-agent swarm finds the cache in the first place. Running black-box on your own perimeter, it fingerprints an exposed inference service that no CVE feed has a plugin for, writes a per-target proof-of-concept with a local model — no exploit pulled from a public repository, no prompt leaving the box — and backtests it in the AI Gym before it ever touches production. It runs on-premise on private AI with a human in the loop for every action that matters, so the one thing it will never do is the thing this CVE punishes: trust an unauthenticated message from the network. For a fuller picture of how continuous, exploit-proven testing differs from a scanner's version check, see our guide to automated penetration testing and the companion analysis of the LiteLLM AI-gateway RCE — the last time the AI-serving stack turned out to be the asset nobody inventoried.
Every vulnerability CISA lists as exploited, with federal due dates: CISA KEV tracker →
Is this exploitable in your environment?
Zero Hunt answers that on your own network: an autonomous AI red team on an on-premise appliance, running on private AI, black-box or gray-box, with a human approving every step that matters. Proof of what is exploitable, the fix, and signed evidence — no data leaves your perimeter.