<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0">
  <channel>
    <title>Most recent entries from all</title>
    <link>https://vulnerability.circl.lu</link>
    <description>Contains only the most 10 recent entries.</description>
    <docs>http://www.rssboard.org/rss-specification</docs>
    <generator>python-feedgen</generator>
    <language>en</language>
    <lastBuildDate>Wed, 07 Oct 2026 09:58:52 +0000</lastBuildDate>
    <item>
      <title>fkie_cve-2026-105755</title>
      <link>https://vulnerability.circl.lu/vuln/fkie_cve-2026-105755</link>
      <description>&lt;p&gt;vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker&amp;#39;s query_key value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim&amp;#39;s identifier can overwrite the cached query embedding so the victim&amp;#39;s documents are scored against the attacker&amp;#39;s query, and shared use counters can also cause a late-interaction cache-miss error. This issue is fixed in version 0.30.0.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker&amp;#39;s query_key value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim&amp;#39;s identifier can overwrite the cached query embedding so the victim&amp;#39;s documents are scored against the attacker&amp;#39;s query, and shared use counters can also cause a late-interaction cache-miss error. This issue is fixed in version 0.30.0.&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://vulnerability.circl.lu/vuln/fkie_cve-2026-105755</guid>
    </item>
    <item>
      <title>GHSA-2phq-3phc-84px — vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id — cross-request integ…</title>
      <link>https://vulnerability.circl.lu/vuln/ghsa-2phq-3phc-84px</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;## Affected&lt;/p&gt;
&lt;p&gt;- **Ecosystem / package:** pip / `vllm`
- **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches.&lt;/p&gt;
&lt;p&gt;## Summary&lt;/p&gt;
&lt;p&gt;On late-interaction `/score` and `/rerank` deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the **caller-controlled** `X-Request-Id` header. A second concurrent request that reuses the victim&amp;#39;s header value replaces the victim&amp;#39;s cached query embedding before document scoring — so the victim&amp;#39;s documents are scored against the **attacker&amp;#39;s** query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim&amp;#39;s `X-Request-Id`. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error.&lt;/p&gt;
&lt;p&gt;This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models.&lt;/p&gt;
&lt;p&gt;## Affected code&lt;/p&gt;
&lt;p&gt;Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491c…&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;## Affected&lt;/p&gt;
&lt;p&gt;- **Ecosystem / package:** pip / `vllm`
- **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches.&lt;/p&gt;
&lt;p&gt;## Summary&lt;/p&gt;
&lt;p&gt;On late-interaction `/score` and `/rerank` deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the **caller-controlled** `X-Request-Id` header. A second concurrent request that reuses the victim&amp;#39;s header value replaces the victim&amp;#39;s cached query embedding before document scoring — so the victim&amp;#39;s documents are scored against the **attacker&amp;#39;s** query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim&amp;#39;s `X-Request-Id`. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error.&lt;/p&gt;
&lt;p&gt;This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models.&lt;/p&gt;
&lt;p&gt;## Affected code&lt;/p&gt;
&lt;p&gt;Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491c…&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://vulnerability.circl.lu/vuln/ghsa-2phq-3phc-84px</guid>
    </item>
  </channel>
</rss>
