GHSA-935W-9G4M-P28P

Vulnerability from github – Published: 2026-10-06 00:02 – Updated: 2026-10-06 00:02
VLAI
Summary
vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle
Details

Affected

  • Ecosystem / package: pip / vllm
  • Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the tool-continuation re-submission has omitted the salt.

Summary

On the GPT-OSS "Harmony" path (POST /v1/responses), a request that uses a built-in or MCP tool runs as a multi-turn loop: after each tool call vLLM re-renders the full next-turn Harmony prompt and re-submits it to the engine. Turn 1 correctly carries request.cache_salt, but the tool-continuation re-submission rebuilds the engine input via tokens_input(token_ids) with no cache_salt. The continuation prefix is therefore cached in the global unsalted namespace even though the caller opted into salting. A second tenant who can guess the low-entropy post-tool history submits the reconstructed continuation (unsalted) and reads exact per-turn cached-token counts from the Responses usage — restoring the prompt-membership oracle that cache_salt is documented to prevent.

Silently dropping a preserved salt after the supported tool workflow is enabled is a broken isolation control: the caller enabled salting and every turn should stay isolated, but continuation turns leak into the shared cache.

This is distinct from GHSA-4qjh-9fv9-r85r (CVE-2025-46570): that advisory is the prefix-cache membership oracle for which cache_salt is the documented mitigation, and its PR-17045 fix does not close this site — the Harmony tool continuation silently drops the preserved salt, caching in the unsalted namespace and leaking exact cached_tokens_per_turn counts from a different sink (the Responses serving continuation, not general TTFT timing).

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

The tool-continuation re-submission rebuilds the engine input with no cache_salt:

# vllm/entrypoints/openai/responses/serving.py Lines 711-715
            if isinstance(context, HarmonyContext):
                token_ids = context.render_for_completion()
                engine_input = tokens_input(token_ids)

                sampling_params.max_tokens = max_model_len - len(token_ids)

Contrast with the correct turn-1 call, which does preserve the caller's salt:

# vllm/entrypoints/openai/responses/serving.py Lines 754-755
        prompt_token_ids = render_for_completion(messages)
        engine_input = tokens_input(prompt_token_ids, cache_salt=request.cache_salt)

tokens_input stores the salt on the engine input only when it is passed, so the continuation input carries none and lands in the unsalted namespace:

# vllm/inputs/engine.py Lines 51-66
def tokens_input(
    prompt_token_ids: list[int],
    *,
    prompt: str | None = None,
    cache_salt: str | None = None,
) -> TokensInput:
    """
    Construct [`TokensInput`][vllm.inputs.engine.TokensInput]
    from optional values.
    """
    inputs = TokensInput(type="token", prompt_token_ids=prompt_token_ids)

    if prompt is not None:
        inputs["prompt"] = prompt
    if cache_salt is not None:
        inputs["cache_salt"] = cache_salt

Impact

An authenticated tenant of a shared deployment can recover whether a guessed post-tool prompt or history was processed by another tenant, with exact cached-token counts rather than noisy latency — the exact prompt-membership oracle cache_salt is documented to prevent. It defeats the multi-user prefix-cache isolation guarantee for salted Harmony tool sessions.

Preconditions: a GPT-OSS Harmony model on /v1/responses; prefix caching enabled (default); an operator-enabled built-in or MCP tool server; the victim sets cache_salt and triggers at least one tool continuation; and the attacker can reconstruct the post-tool history closely enough to match the token prefix. The AC:H metric reflects that guessable-history precondition.

Suggested Fix

Propagate request.cache_salt into every Harmony (and Parsable) tool-continuation re-submission — at the continuation call site call tokens_input(token_ids, cache_salt=request.cache_salt), mirroring the correct turn-1 call. Carry the salt on the HarmonyContext (thread the originating request into the context) so no continuation path can omit it:

# vllm/entrypoints/openai/responses/serving.py
             if isinstance(context, HarmonyContext):
                 token_ids = context.render_for_completion()
-                engine_input = tokens_input(token_ids)
+                engine_input = tokens_input(
+                    token_ids,
+                    cache_salt=(
+                        context.request.cache_salt
+                        if context.request is not None
+                        else None
+                    ),
+                )

with HarmonyContext.__init__ gaining a request: ResponsesRequest | None = None parameter (stored as self.request) that _create_responses passes when constructing the context. The continuation prefix is then cached in the victim's salted namespace, mirroring turn 1.

Suggested regression test: assert cached_tokens_per_turn == 0 for a different-salt probe against a salted victim continuation (the four-way control from the proof of concept).

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.


Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51818

Show details on source website

{
  "affected": [
    {
      "package": {
        "ecosystem": "PyPI",
        "name": "vllm"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "0"
            },
            {
              "fixed": "0.30.0"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-105752"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-200",
      "CWE-524"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-10-06T00:02:02Z",
    "nvd_published_at": null,
    "severity": "LOW"
  },
  "details": "## Affected\n\n- **Ecosystem / package:** pip / `vllm`\n- **Affected versions:** vLLM \u2264 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the tool-continuation re-submission has omitted the salt.\n\n## Summary\n\nOn the GPT-OSS \"Harmony\" path (`POST /v1/responses`), a request that uses a built-in or MCP tool runs as a multi-turn loop: after each tool call vLLM re-renders the full next-turn Harmony prompt and re-submits it to the engine. Turn 1 correctly carries `request.cache_salt`, but the tool-continuation re-submission rebuilds the engine input via `tokens_input(token_ids)` with **no** `cache_salt`. The continuation prefix is therefore cached in the global *unsalted* namespace even though the caller opted into salting. A second tenant who can guess the low-entropy post-tool history submits the reconstructed continuation (unsalted) and reads exact per-turn cached-token counts from the Responses usage \u2014 restoring the prompt-membership oracle that `cache_salt` is documented to prevent.\n\nSilently dropping a preserved salt *after* the supported tool workflow is enabled is a broken isolation control: the caller enabled salting and every turn should stay isolated, but continuation turns leak into the shared cache.\n\nThis is distinct from [GHSA-4qjh-9fv9-r85r](https://github.com/vllm-project/vllm/security/advisories/GHSA-4qjh-9fv9-r85r) ([CVE-2025-46570](https://nvd.nist.gov/vuln/detail/CVE-2025-46570)): that advisory is the prefix-cache membership oracle for which `cache_salt` is the documented mitigation, and its PR-17045 fix does not close this site \u2014 the Harmony tool continuation silently drops the preserved salt, caching in the unsalted namespace and leaking exact `cached_tokens_per_turn` counts from a different sink (the Responses serving continuation, not general TTFT timing).\n\n## Affected code\n\nLinks pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1):\n\n- **The drop (sink):** [`vllm/entrypoints/openai/responses/serving.py#L712-L713`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/serving.py#L712-L713) \u2014 `token_ids = context.render_for_completion()` then `engine_input = tokens_input(token_ids)`, with no `cache_salt`.\n- **Correct turn-1 call for contrast:** [`vllm/entrypoints/openai/responses/serving.py#L755`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/serving.py#L755) \u2014 `tokens_input(prompt_token_ids, cache_salt=request.cache_salt)`.\n- **`tokens_input` stores the salt only if passed:** [`vllm/inputs/engine.py#L51-L66`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/inputs/engine.py#L51-L66) (`if cache_salt is not None: inputs[\"cache_salt\"] = cache_salt`).\n- **The engine request copies only the current input\u0027s salt:** [`vllm/v1/engine/input_processor.py#L380`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/engine/input_processor.py#L380) (`cache_salt=decoder_inputs.get(\"cache_salt\")` \u2192 `None` for the continuation).\n- **Prefix-cache hashing keys on the salt only when present:** [`vllm/v1/core/kv_cache_utils.py#L560-L561`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/v1/core/kv_cache_utils.py#L560-L561) (`[request.cache_salt] if (start_token_idx == 0 and request.cache_salt) else []`).\n- **The oracle the attacker reads:** [`vllm/entrypoints/openai/responses/serving.py#L909`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/serving.py#L909) (`cached_tokens_per_turn`).\n- **The documented control being defeated:** [`vllm/entrypoints/openai/responses/protocol.py#L235`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/responses/protocol.py#L235) (`cache_salt` field).\n\nThe tool-continuation re-submission rebuilds the engine input with no `cache_salt`:\n\n```python\n# vllm/entrypoints/openai/responses/serving.py Lines 711-715\n            if isinstance(context, HarmonyContext):\n                token_ids = context.render_for_completion()\n                engine_input = tokens_input(token_ids)\n\n                sampling_params.max_tokens = max_model_len - len(token_ids)\n```\n\nContrast with the correct turn-1 call, which does preserve the caller\u0027s salt:\n\n```python\n# vllm/entrypoints/openai/responses/serving.py Lines 754-755\n        prompt_token_ids = render_for_completion(messages)\n        engine_input = tokens_input(prompt_token_ids, cache_salt=request.cache_salt)\n```\n\n`tokens_input` stores the salt on the engine input only when it is passed, so the continuation input carries none and lands in the unsalted namespace:\n\n```python\n# vllm/inputs/engine.py Lines 51-66\ndef tokens_input(\n    prompt_token_ids: list[int],\n    *,\n    prompt: str | None = None,\n    cache_salt: str | None = None,\n) -\u003e TokensInput:\n    \"\"\"\n    Construct [`TokensInput`][vllm.inputs.engine.TokensInput]\n    from optional values.\n    \"\"\"\n    inputs = TokensInput(type=\"token\", prompt_token_ids=prompt_token_ids)\n\n    if prompt is not None:\n        inputs[\"prompt\"] = prompt\n    if cache_salt is not None:\n        inputs[\"cache_salt\"] = cache_salt\n```\n\n## Impact\n\nAn authenticated tenant of a shared deployment can recover whether a guessed post-tool prompt or history was processed by another tenant, with exact cached-token counts rather than noisy latency \u2014 the exact prompt-membership oracle `cache_salt` is documented to prevent. It defeats the multi-user prefix-cache isolation guarantee for salted Harmony tool sessions.\n\nPreconditions: a GPT-OSS Harmony model on `/v1/responses`; prefix caching enabled (default); an operator-enabled built-in or MCP tool server; the victim sets `cache_salt` and triggers at least one tool continuation; and the attacker can reconstruct the post-tool history closely enough to match the token prefix. The `AC:H` metric reflects that guessable-history precondition.\n\n\n## Suggested Fix\n\nPropagate `request.cache_salt` into every Harmony (and Parsable) tool-continuation re-submission \u2014 at the continuation call site call `tokens_input(token_ids, cache_salt=request.cache_salt)`, mirroring the correct turn-1 call. Carry the salt on the `HarmonyContext` (thread the originating `request` into the context) so no continuation path can omit it:\n\n```diff\n# vllm/entrypoints/openai/responses/serving.py\n             if isinstance(context, HarmonyContext):\n                 token_ids = context.render_for_completion()\n-                engine_input = tokens_input(token_ids)\n+                engine_input = tokens_input(\n+                    token_ids,\n+                    cache_salt=(\n+                        context.request.cache_salt\n+                        if context.request is not None\n+                        else None\n+                    ),\n+                )\n```\n\nwith `HarmonyContext.__init__` gaining a `request: ResponsesRequest | None = None` parameter (stored as `self.request`) that `_create_responses` passes when constructing the context. The continuation prefix is then cached in the victim\u0027s salted namespace, mirroring turn 1.\n\nSuggested regression test: assert `cached_tokens_per_turn == 0` for a different-salt probe against a salted victim continuation (the four-way control from the proof of concept).\n\n## Credit\n\n**Reported by:** Patch the Planet (Trail of Bits + OpenAI collaboration)\n\nThis vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.\n\n---\n\n**Proposed fix:** a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51818",
  "id": "GHSA-935w-9g4m-p28p",
  "modified": "2026-10-06T00:02:03Z",
  "published": "2026-10-06T00:02:02Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/security/advisories/GHSA-935w-9g4m-p28p"
    },
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/pull/50195"
    },
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/pull/51818"
    },
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/commit/6a2a2bb02b563b83f946012959fd3927984d072a"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/vllm-project/vllm"
    },
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/releases/tag/v0.30.0"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:N/I:L/A:N",
      "type": "CVSS_V3"
    }
  ],
  "summary": "vLLM: Harmony tool continuations drop `cache_salt` \u2014 restoring a cross-tenant prefix-cache membership oracle"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Loading…

Loading…

Related by attack behaviour

Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.


Loading…