Skip to main content
QUIETLYTIC
Vulnerability

10 CVEs: vLLM & vLLM Hardware Plugin for Intel Gaudi (CVE-2026-73560)

Ten 2026 vLLM CVEs assessed against NVD, GitHub Advisory, and CVE.org data: an SSRF and file-read flaw, a cross-user data leak, and an OpenAI API auth bypass.

CVE-2026-73560
Threat Level
MEDIUM
CVSS
6.5
Status
Monitored
Confidence
Medium
Affected Products
vLLM, vLLM Hardware Plugin for Intel Gaudi

Full CVE Roster

All 10 CVEs from this release, ready to paste into a tracker, ticket, or SIEM search.

CVE ID Title CVSS Severity KEV
CVE-2026-73560 — 6.5 medium
CVE-2026-55646 — 6.5 medium
CVE-2026-73558 — 5.3 medium
CVE-2026-73555 — 5.3 medium
CVE-2026-73556 — 5.3 medium
CVE-2026-71486 — 4.3 medium
CVE-2026-73557 — 0.0 medium
CVE-2026-48746 — — —
CVE-2026-27765 — — —
CVE-2026-92365 — — —

vLLM, the open-source inference and serving engine behind much of the self-hosted large-language-model deployment stack, disclosed ten CVEs between June and September 2026. The most severe, an SSRF and arbitrary local file read in its multimodal media loader (CVE-2026-73560, CVSS 6.5), sits alongside a tied-severity memory-exhaustion denial of service, a cross-user data leak inside a GPU kernel, an OpenAI-compatible API authentication bypass, and six further denial-of-service and information-disclosure issues. None are listed in CISA’s Known Exploited Vulnerabilities catalog as of this writing, and each CVE’s evidence traces to exactly one named source in our ledger — NVD, GitHub Advisory Database, or CVE.org — so confidence across this roundup is medium throughout, not high.

CVE-2026-73560: SSRF and local file read in the MiMoV2Omni multimodal processor

Affects vLLM before 0.26.0. Per the GitHub Advisory Database record, the multimodal processor for MiMoV2Omni-based models fetches caller-supplied image and audio URLs and opens caller-supplied local file paths directly, bypassing the SSRF and allowed-local-media-path checks vLLM’s shared media-connector utility enforces on its other multimodal code paths. That gap lets a request make the server reach internal or local resources under the server’s own network and filesystem access. CWE-918 (SSRF). CVSS base 6.5 (medium; vector AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N). Fixed in 0.26.0. Confidence: medium.

CVE-2026-55646: unbounded memory use in audio transcription and translation uploads

Affects vLLM 0.22.0 through 0.23.0. Per the NVD record, the /v1/audio/transcriptions and /v1/audio/translations endpoints read an entire uploaded audio file into memory before checking it against the documented VLLM_MAX_AUDIO_CLIP_FILESIZE_MB limit (default 25 MB) — the size check runs too late to prevent the memory allocation it’s meant to bound. CWE-400 / CWE-770. CVSS base 6.5 (medium; vector AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H). Confidence: medium — NVD only; this ledger entry carries no fixed-version data, a gap we’re noting rather than guessing at.

CVE-2026-73558: cross-user data leak via integer overflow in the act_and_mul_kernel

Affects vLLM before 0.26.0 (pip:vllm). Per the GitHub Advisory Database record, an integer overflow in the act_and_mul_kernel GPU kernel can, under certain batching conditions, let part of one user’s inference output reach the response returned to a different user in the same inference batch — the record describes this as ranging up to a full copy of the first user’s result reaching a later request. CWE-190 (Integer Overflow). CVSS base 5.3 (medium; vector AV:N/AC:H/PR:N/UI:R/S:U/C:H/I:N/A:N — network attack vector offset by high attack complexity and a required user-interaction condition). Fixed in 0.27.0. Confidence: medium.

CVE-2026-73555: internal path and username disclosure via validation error messages

Affects vLLM before 0.26.0. Per the GitHub Advisory Database record, a malformed API request triggers a Pydantic RequestValidationError, and vLLM’s own exception handler converts that error to a string and returns it unredacted — exposing the server’s internal file paths and the OS username running vLLM to any unauthenticated caller who sends a malformed request. CWE-209 (Information Exposure Through an Error Message). CVSS base 5.3 (medium; vector AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N). Fixed in 0.26.0. Confidence: medium.

CVE-2026-73556: ReDoS in the lm-format-enforcer structured-output backend

Affects vLLM before 0.26.0. Per the GitHub Advisory Database record, an earlier fix for a separate ReDoS advisory (GHSA-rwxx-mrjm-wc2m) added a compile timeout to vLLM’s xgrammar and outlines structured-output backends — but missed the lm-format-enforcer backend, which still compiles caller-supplied structured_outputs.regex patterns without a timeout, leaving that one backend exposed to the same denial-of-service class the earlier fix was meant to close. CWE-400 (Uncontrolled Resource Consumption) and CWE-1333 (ReDoS). CVSS base 5.3 (medium; vector AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L). Fixed in 0.26.0. Confidence: medium.

CVE-2026-71486: unbounded output processing in the derender endpoints

Affects vLLM before 0.26.0. Per the GitHub Advisory Database record, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects and postprocess every nested token-ID list directly, without the model context-length, max-tokens, max-sequence-count, or choice-count limits vLLM’s normal render/generate path enforces. That lets a caller drive resource consumption the standard endpoints are designed to cap. CWE-400 / CWE-770 (uncontrolled resource consumption / allocation without limits). CVSS base 4.3 (medium; vector AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L). Fixed in 0.26.0. Confidence: medium.

CVE-2026-73557: incomplete fix for a 2025 prompt-embedding race condition

Affects the pre-0.26.0 vLLM revision the GitHub Advisory record cites. Per that record, the follow-up protection shipped for CVE-2025-62164 wraps serialized prompt-embedding reconstruction in PyTorch’s check_sparse_tensor_invariants() context manager, but PyTorch 2.11.0 implements that context with save/restore operations a concurrent second prompt part can interleave with — letting the original issue resurface under concurrent requests. CWE-362 (concurrent execution using shared resource with improper synchronization). This is a data quirk worth stating plainly: GitHub Advisory Database records this entry’s numeric CVSS base score as 0 while separately labeling its qualitative severity “medium” — we’re reporting that inconsistency as it appears in the source rather than resolving it in either direction. Fixed in 0.26.0. Confidence: medium.

CVE-2026-48746: OpenAI-compatible API authentication bypass

Affects vLLM 0.3.0 through 0.22.0. Per the CVE.org record, vLLM’s OpenAI-compatible API relies on Starlette/ASGI middleware to enforce its AuthenticationMiddleware, and a trust assumption between that middleware and the underlying ASGI server let requests reach protected endpoints without supplying the configured VLLM_API_KEY or --api-key value. The record’s reference list includes an independent x41 D-Sec advisory covering the underlying Starlette trust issue. CWE-444 (Inconsistent Interpretation of HTTP Requests). No CVSS score is recorded in this ledger entry. Fixed in 0.22.0. Confidence: medium — CVE.org only, no independent NVD or GitHub Advisory corroboration in our data.

CVE-2026-27765: denial of service in the Intel Gaudi hardware plugin

Affects vLLM’s Hardware Plugin for Intel Gaudi, versions before 0.16.0 — a separate package from vLLM’s core PyPI distribution, relevant only to deployments running vLLM on Intel Gaudi accelerators. Per the NVD record, which cites an underlying Intel security advisory, improper input validation in a Ring 3 user-application component can let an authenticated, authorized local user with a low-complexity attack trigger a denial of service. NVD’s own description hedges this as something that “may potentially occur,” which we read as the source itself carrying some uncertainty about the exact trigger conditions rather than a confirmed, unconditional impact. CWE-20 (Improper Input Validation). No CVSS score is recorded in this ledger entry. Confidence: medium.

CVE-2026-92365: algorithmic complexity issue in thinking-budget state tracking

Affects vLLM up to 0.29.0. Per the CVE.org record, unspecified functionality in vllm/v1/sample/thinking_budget_state.py — the module tracking per-request “thinking budget” state for reasoning models — has inefficient algorithmic complexity that the record states can be triggered remotely. As of this writing, the record itself notes the fix is an open pull request that hasn’t merged yet, and it gives no further detail on the specific trigger. No CVSS score is recorded in this ledger entry, and this is the least-detailed record in the batch. Confidence: medium, with an explicit gap — the evidence doesn’t support characterizing real-world impact beyond “inefficient algorithmic complexity,” and the fix isn’t final.

Confidence and evidence gaps

Every CVE in this roundup traces to exactly one source in our ledger (NVD, GitHub Advisory Database, or CVE.org), so confidence is medium across the board under our standard model — two or more independent sources agreeing would raise a finding to high, and none here reach that bar. Specific gaps worth naming rather than papering over: no EPSS score is present in our data for any of these ten CVEs; the field-provenance table records no per-field source URL, so citations above are attributed by source name rather than a direct link; CVE-2026-55646 has no fixed-version data in our ledger; CVE-2026-73557’s numeric CVSS score and its qualitative severity label disagree within the same source record; and CVE-2026-92365’s fix was still an open, unmerged pull request as of this writing.

Why this matters

vLLM sits underneath a large share of self-hosted LLM inference deployments, which is what makes an SSRF in its multimodal ingestion path and a cross-user data leak inside its serving path more consequential than their medium CVSS scores alone suggest — a request reaching internal or local resources through an image/audio loader, or one tenant’s output surfacing in another tenant’s response, cuts against guarantees a multi-user inference service is expected to hold. The OpenAI-compatible API authentication bypass adds a third dimension: an unauthenticated request reaching an endpoint meant to require a key. None of these ten are confirmed under active exploitation, and most require narrow preconditions (fixed dependency versions, specific batching conditions, or hardware-plugin deployments) — but the volume and diversity of issues across a single release cycle argues for staying current with vLLM’s patch releases rather than treating any one of these in isolation.

Frequently Asked Questions

Are any of these vLLM vulnerabilities being actively exploited? No. None of the ten are listed in CISA’s Known Exploited Vulnerabilities catalog as of this writing, and our source data contains no exploitation reports for any of them.

Which vLLM version fixes these issues? Most of the CVEs with a documented fix point to 0.26.0 or 0.27.0. CVE-2026-48746 (the auth bypass) is fixed in 0.22.0. Two entries — CVE-2026-55646 and CVE-2026-92365 — have no fixed-version data in our ledger; for those, check the linked GitHub Advisory Database or CVE.org record directly before assuming a specific release resolves them.

Does the vLLM authentication bypass affect deployments that never set an API key? The CVE describes a bypass of the AuthenticationMiddleware layer itself, which only matters where VLLM_API_KEY or --api-key was configured in the first place — a deployment that never enabled API-key auth has no auth layer for this bug to bypass, though it also has no auth layer protecting it to begin with.

Is the cross-user data leak (CVE-2026-73558) a prompt-injection or model-level issue? No. It’s an integer overflow in a GPU kernel’s batch-handling code (CWE-190), not a property of the model’s own outputs or a prompt-injection vector — the leak is at the serving-infrastructure layer, between requests in the same batch.


Data sourced from National Vulnerability Database (NVD), GitHub Advisory Database, and CVE.org records, evaluated September 2026. This product uses the NVD API but is not endorsed or certified by the NVD. See more vulnerability intelligence.

Report an error

Found a factual error, an outdated figure, or a broken source link? Let us know and our editorial desk will review it.


Sources & evidence

01 National Vulnerability Database (NVD)
02 GitHub Advisory Database
03 CVE.org

Related intelligence


Analyst tools