CVE-2026-105755 (GCVE-0-2026-105755)
Vulnerability from cvelistv5 – Published: 2026-10-05 22:47 – Updated: 2026-10-05 22:47
VLAI
EPSS
VEX
Title
vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id — cross-request integrity break and induced errors on `/score` and `/rerank`
Summary
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker's query_key value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim's identifier can overwrite the cached query embedding so the victim's documents are scored against the attacker's query, and shared use counters can also cause a late-interaction cache-miss error. This issue is fixed in version 0.30.0.
Severity
4.2 (Medium)
CWE
- CWE-639 - Authorization Bypass Through User-Controlled Key
Assigner
References
4 references
| URL | Tags |
|---|---|
| https://github.com/vllm-project/vllm/security/adv… | x_refsource_CONFIRM |
| https://github.com/vllm-project/vllm/pull/51445 | x_refsource_MISC |
| https://github.com/vllm-project/vllm/commit/ee17d… | x_refsource_MISC |
| https://github.com/vllm-project/vllm/releases/tag… | x_refsource_MISC |
Impacted products
1 product
| Vendor | Product | Version | CPE status | |
|---|---|---|---|---|
| vllm-project | vllm |
Affected:
< 030.0
|
guessed |
{
"containers": {
"cna": {
"affected": [
{
"product": "vllm",
"vendor": "vllm-project",
"versions": [
{
"status": "affected",
"version": "\u003c 030.0"
}
]
}
],
"descriptions": [
{
"lang": "en",
"value": "vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker\u0027s query_key value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim\u0027s identifier can overwrite the cached query embedding so the victim\u0027s documents are scored against the attacker\u0027s query, and shared use counters can also cause a late-interaction cache-miss error. This issue is fixed in version 0.30.0."
}
],
"metrics": [
{
"cvssV3_1": {
"attackComplexity": "HIGH",
"attackVector": "NETWORK",
"availabilityImpact": "LOW",
"baseScore": 4.2,
"baseSeverity": "MEDIUM",
"confidentialityImpact": "NONE",
"integrityImpact": "LOW",
"privilegesRequired": "LOW",
"scope": "UNCHANGED",
"userInteraction": "NONE",
"vectorString": "CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:N/I:L/A:L",
"version": "3.1"
}
}
],
"problemTypes": [
{
"descriptions": [
{
"cweId": "CWE-639",
"description": "CWE-639: Authorization Bypass Through User-Controlled Key",
"lang": "en",
"type": "CWE"
}
]
}
],
"providerMetadata": {
"dateUpdated": "2026-10-05T22:47:54.854Z",
"orgId": "a0819718-46f1-4df5-94e2-005712e83aaa",
"shortName": "GitHub_M"
},
"references": [
{
"name": "https://github.com/vllm-project/vllm/security/advisories/GHSA-2phq-3phc-84px",
"tags": [
"x_refsource_CONFIRM"
],
"url": "https://github.com/vllm-project/vllm/security/advisories/GHSA-2phq-3phc-84px"
},
{
"name": "https://github.com/vllm-project/vllm/pull/51445",
"tags": [
"x_refsource_MISC"
],
"url": "https://github.com/vllm-project/vllm/pull/51445"
},
{
"name": "https://github.com/vllm-project/vllm/commit/ee17d0d869203ef9a35ad73358a4987bba14b1fc",
"tags": [
"x_refsource_MISC"
],
"url": "https://github.com/vllm-project/vllm/commit/ee17d0d869203ef9a35ad73358a4987bba14b1fc"
},
{
"name": "https://github.com/vllm-project/vllm/releases/tag/v0.30.0",
"tags": [
"x_refsource_MISC"
],
"url": "https://github.com/vllm-project/vllm/releases/tag/v0.30.0"
}
],
"source": {
"advisory": "GHSA-2phq-3phc-84px",
"discovery": "UNKNOWN"
},
"title": "vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id \u2014 cross-request integrity break and induced errors on `/score` and `/rerank`"
}
},
"cveMetadata": {
"assignerOrgId": "a0819718-46f1-4df5-94e2-005712e83aaa",
"assignerShortName": "GitHub_M",
"cveId": "CVE-2026-105755",
"datePublished": "2026-10-05T22:47:54.854Z",
"dateReserved": "2026-10-05T19:11:07.947Z",
"dateUpdated": "2026-10-05T22:47:54.854Z",
"state": "PUBLISHED"
},
"dataType": "CVE_RECORD",
"dataVersion": "5.2",
"vulnerability-lookup:meta": {
"epss": {
"cve": "CVE-2026-105755",
"date": "2026-10-06",
"epss": "0.00204",
"percentile": "0.0943"
},
"nvd": {
"cve": {
"affected": [
{
"affectedData": [
{
"product": "vllm",
"vendor": "vllm-project",
"versions": [
{
"status": "affected",
"version": "\u003c 030.0"
}
]
}
],
"source": "security-advisories@github.com"
}
],
"cveTags": [],
"descriptions": [
{
"lang": "en",
"value": "vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker\u0027s query_key value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim\u0027s identifier can overwrite the cached query embedding so the victim\u0027s documents are scored against the attacker\u0027s query, and shared use counters can also cause a late-interaction cache-miss error. This issue is fixed in version 0.30.0."
}
],
"id": "CVE-2026-105755",
"lastModified": "2026-10-06T14:59:48.280",
"metrics": {
"cvssMetricV31": [
{
"cvssData": {
"attackComplexity": "HIGH",
"attackVector": "NETWORK",
"availabilityImpact": "LOW",
"baseScore": 4.2,
"baseSeverity": "MEDIUM",
"confidentialityImpact": "NONE",
"integrityImpact": "LOW",
"privilegesRequired": "LOW",
"scope": "UNCHANGED",
"userInteraction": "NONE",
"vectorString": "CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:N/I:L/A:L",
"version": "3.1"
},
"exploitabilityScore": 1.6,
"impactScore": 2.5,
"source": "security-advisories@github.com",
"type": "Secondary"
}
]
},
"published": "2026-10-05T23:17:02.167",
"references": [
{
"source": "security-advisories@github.com",
"url": "https://github.com/vllm-project/vllm/commit/ee17d0d869203ef9a35ad73358a4987bba14b1fc"
},
{
"source": "security-advisories@github.com",
"url": "https://github.com/vllm-project/vllm/pull/51445"
},
{
"source": "security-advisories@github.com",
"url": "https://github.com/vllm-project/vllm/releases/tag/v0.30.0"
},
{
"source": "security-advisories@github.com",
"url": "https://github.com/vllm-project/vllm/security/advisories/GHSA-2phq-3phc-84px"
}
],
"sourceIdentifier": "security-advisories@github.com",
"vulnStatus": "Undergoing Analysis",
"weaknesses": [
{
"description": [
{
"lang": "en",
"value": "CWE-639"
}
],
"source": "security-advisories@github.com",
"type": "Primary"
}
]
}
},
"redhat_vex": {
"aggregate_severity": "Moderate",
"current_release_date": "2026-10-06T00:11:33+00:00",
"cve": "CVE-2026-105755",
"id": "CVE-2026-105755",
"initial_release_date": "2026-10-05T22:47:54.854000+00:00",
"product_status:known_affected": "27",
"source": "Red Hat CSAF VEX",
"status": "final",
"title": "vllm: vllm: Scoring result corruption via user-controlled request identifier",
"url": "https://security.access.redhat.com/data/csaf/v2/vex/2026/cve-2026-105755.json",
"version": "3"
}
}
}
Loading…
Loading…
Experimental. This forecast is provided for visualization only and may change without notice. Do not use it for operational decisions.
Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.
Loading…
Loading…
The MITRE ATT&CK techniques below are AI-generated suggestions, inferred from the description of the
vulnerability by the CIRCL/vulnerability-attack-technique-classification-roberta-base
model, served locally by ML-Gateway.
They have not been verified by an analyst and are provided for guidance only.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.
Browse all ATT&CK techniques and the vulnerabilities related to each.
The approach is described in our paper Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion.
Browse all ATT&CK techniques and the vulnerabilities related to each.
Loading…
Related by attack behaviour
Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.
Loading…