Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

vLLM cache_salt flaw lets one request crash EngineCore on LMCache-MP deployments

CVE-2026-105756 lets a single API request with a malformed cache_salt crash vLLM's EngineCore on LMCache-MP deployments. Patched in version 0.30.0.

A medium-severity vulnerability in vLLM allows a single malicious API request to crash the serving engine for all concurrent users on deployments running the LMCache-MP KV connector, according to a GitHub security advisory1.

The advisory, tracked as CVE-2026-105756 and GHSA-2823-qmq8-rwvj, said vLLM's OpenAI-compatible request models for Completions, Chat Completions, and Responses accept a client-supplied `cache_salt` field but validate it only as a non-empty string, with no character or length restrictions. On deployments where the built-in LMCache-MP KV connector is enabled, that value is forwarded unguarded into the scheduler's per-step cache lookup.

LMCache enforces a stricter check inside IPCCacheServerKey's `__post_init__` method, which raises a `ValueError` for any `cache_salt` containing `@`, `/`, `\`, or NUL characters, or exceeding 128 characters. Neither the LMCache-MP connector lookup call site nor Scheduler.schedule() wraps that call in a request-scoped exception handler, so the `ValueError` propagates uncaught into EngineCore's top-level loop. The advisory said a request as simple as `cache_salt="/"` is sufficient to take down the engine for every user sharing it.

The advisory classified the issue as a denial-of-service flaw affecting availability only. Versions through vLLM 0.25.1 are confirmed affected; the advisory said the lower bound may extend further back, depending on how long the loose validator and unguarded scheduling path have been present.

A fix is available in vLLM 0.30.0, according to the advisory, which also pointed to public pull request 51444 as the proposed remediation. The flaw applies only when the LMCache-MP KV connector is enabled; deployments not using that connector are not exposed.

This is the third medium-severity denial-of-service advisory disclosed against vLLM in recent days. A separate advisory covered CVE-2026-105760, a CPU and memory exhaustion bug in the GLMGA video sampler affecting versions 0.23.0rc2 through 0.29.x[1]. Another, CVE-2026-105758, described an unauthenticated memory-exhaustion path in Qwen2-VL and Qwen3-VL video backends[2]. ANALYSIS The cluster of disclosures across distinct vLLM subsystems — video sampling, multimodal backends, and now KV-cache connectors — widens the attack surface that operators of public-facing vLLM endpoints must audit when planning upgrades to 0.30.0.