Swati KhandelwalOct 07, 2026Vulnerability / Artificial Intelligence
A critical vulnerability in LMCache, open-source software that speeds up large language model (LLM) servers such as vLLM, lets an attacker run code on the cache server without logging in, and no fixed version is available.
The flaw is in LMCache’s multiprocess mode, where the cache runs as a standalone server that LLM workers reach over the ZeroMQ messaging library. A single network message to that server can run commands as the user the LMCache process runs as.
The server can be reached from another machine only when an operator sets it to listen on a routable address, rather than the localhost it uses by default.
JFrog disclosed the flaw on October 7 and assigned it a severity score of 9.8 out of 10, in the critical range, the rating it gives a server bound to a routable address.
The vulnerability, tracked as CVE-2026-105192, affects LMCache from version 0.3.9, released in October 2025, through 0.5.5, the latest stable release, and is also present in the 0.5.6 release candidates and the development branch. No fixed version exists.
Whether a server is exposed comes down to one setting. By default, the multiprocess server listens only on the local machine, so another host cannot reach it. It becomes reachable when an operator starts it with a routable address, much like multi-node deployments share a cache across machines.
LMCache’s own example Kubernetes deployment starts the server that way, listening on every network interface. A copy of LMCache running inside a single vLLM process does not open the port at all.
The ZeroMQ socket the multiprocess server opens for worker processes to register and share cached data has no authentication. One type of message is unpacked with pickle, a Python format that can carry code and run it as the data is decoded. The server unpacks it while still reading the message’s arguments, before any check of the message’s type, so a crafted message can run the sender’s code.
The code runs with the privileges of the LMCache process. On the project’s official container images, that process runs as root, according to JFrog. The flaw was found by Yuval Moravchick of JFrog’s security research team.
There is no patched release. Until one ships, JFrog advises operators not to assign the multiprocess server a routable address and to keep its port on the local machine or on a trusted cluster network. A firewall that limits who can reach the port lowers the risk but does not remove it, because any host that can still open a connection can run code.
LMCache has not published a security advisory for the flaw. JFrog’s advisory does not provide operators with a way to determine whether a server has already been attacked.
Other Reports and a Related vLLM Fix
Separately, a GitHub user opened six additional LMCache security reports on October 6, the day before CVE-2026-105192 was made public. They allege unauthenticated access to cached data belonging to different tenants, as well as to several network services that execute commands without a login.
The reports come from one account, rest on proof-of-concept claims, and have no CVE, no confirmation from the maintainers, and no fix. One points to a default LMCache that has since changed: an admin HTTP server that listened on every network interface in 0.5.5 listens only on the local host in the 0.5.6 release candidates.
A related flaw in vLLM is already fixed. Before version 0.30.0, released September 22, a single request carrying a malformed cache_salt value could crash the engine on deployments that use the LMCache multiprocess connector, a denial-of-service bug tracked as CVE-2026-105756. It is rated 6.5 and does not allow code execution.
The core mistake, handing data from an unauthenticated network socket to pickle, is the same one researchers found across other AI inference frameworks in November 2025, in a group of flaws they called ShadowMQ. Whether LMCache’s code shares a common source with those projects has not been established.
Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.

