CVE-2026-54234
Last modified
CVE-2026-54234 is a high-severity vulnerability rated 7.5/10 on the CVSS scale. vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the model vocabulary size boundary value, which is then converted to negative one when the engine selects the next live token for a request and is written back into the drafter's input ids; that out-of-vocabulary value is later consumed by the model's embedding and attention path and crashes the engine worker with a GPU device-side assertion. EPSS estimates a 0.34% chance of exploitation in the next 30 days.
Description
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the model vocabulary size boundary value, which is then converted to negative one when the engine selects the next live token for a request and is written back into the drafter's input ids; that out-of-vocabulary value is later consumed by the model's embedding and attention path and crashes the engine worker with a GPU device-side assertion. The same triggering request sequence is reachable through the public gRPC Generate and Abort endpoints, so a remote client that can send generation requests can crash the shared engine worker, aborting concurrent requests and causing a service-wide denial of service for other clients of the deployment until the worker is restarted. This issue is fixed in version 0.24.0.
Metrics
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
Weakness Enumeration
Affected Software
| Vendor | Product | Versions |
|---|---|---|
| Vllm | Vllm | < 0.24.0 |
References
- https://github.com/vllm-project/vllm/pull/44744Issue Tracking, Patch
- https://github.com/vllm-project/vllm/security/advisories/GHSA-8wr5-jm2h-8r4fVendor Advisory, Exploit
- https://github.com/vllm-project/vllm/security/advisories/GHSA-8wr5-jm2h-8r4fVendor Advisory, Exploit
Timeline
- Published
- Last Modified
- Status
- Analyzed
Frequently Asked Questions
What is CVE-2026-54234?
How severe is CVE-2026-54234?
How do I fix CVE-2026-54234?
How Strix Helps
- How Strix found a critical auth bypass in etcdStrix autonomously discovered a critical authentication bypass in etcd, later designated CVE-2026-33413.
- Autonomous PentestingAI agents that find and validate exploitable vulnerabilities like this one across your applications.
- PR ReviewsPentest every pull request so vulnerable code is caught before it ships to production.
- AI Penetration TestingHow AI-driven penetration testing continuously covers your attack surface.
Related CVEs from 2026
- CVE-2026-54229A race condition was found in the abrt-dbus D-Bus service's …7
- CVE-2026-5423@neo4j/graphql library versions prior to 7.5.6 fail to verif…8.2
- CVE-2026-54230A symlink following vulnerability was found in the ABRT post…7.8
- CVE-2026-54231A content injection vulnerability was found in the ABRT post…5.5
- CVE-2026-54232vLLM is an inference and serving engine for large language m…8.8
- CVE-2026-54233vLLM is an inference and serving engine for large language m…6.5
- CVE-2026-54235vLLM is an inference and serving engine for large language m…6.5
- CVE-2026-54236vLLM is an inference and serving engine for large language m…5.3
- CVE-2026-54242Statamic is a Laravel and Git powered content management sys…4.9
- CVE-2026-54243Statamic is a Laravel and Git powered content management sys…6.1
- CVE-2026-54244Statamic is a Laravel and Git powered content management sys…3.5
- CVE-2026-54249Pydantic AI is a Python agent framework for building Generat…6.8
Are you affected by CVE-2026-54234?
Run a free Strix scan to check your systems for this vulnerability.
Scan your code nowSource: NVD / NIST
