<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://vulnerability.circl.lu/rss/recent/all/10</id>
  <title>Most recent entries from all</title>
  <updated>2026-10-07T22:25:23.846603+00:00</updated>
  <author>
    <name>Vulnerability-Lookup</name>
    <email>info@circl.lu</email>
  </author>
  <link href="https://vulnerability.circl.lu" rel="alternate"/>
  <generator uri="https://lkiesow.github.io/python-feedgen" version="1.0.0">python-feedgen</generator>
  <subtitle>Contains only the most 10 recent entries.</subtitle>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/fkie_cve-2026-105755</id>
    <title>fkie_cve-2026-105755</title>
    <updated>2026-10-07T22:25:23.860401+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml">
        <p>vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker's query_key value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim's identifier can overwrite the cached query embedding so the victim's documents are scored against the attacker's query, and shared use counters can also cause a late-interaction cache-miss error. This issue is fixed in version 0.30.0.</p>
      </div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/fkie_cve-2026-105755"/>
  </entry>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/ghsa-2phq-3phc-84px</id>
    <title>GHSA-2phq-3phc-84px — vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id — cross-request integ…</title>
    <updated>2026-10-07T22:25:23.860566+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> PyPI: vllm</p>
<p>## Affected</p>
<p>- **Ecosystem / package:** pip / `vllm`
- **Affected versions:** vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches.</p>
<p>## Summary</p>
<p>On late-interaction `/score` and `/rerank` deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the **caller-controlled** `X-Request-Id` header. A second concurrent request that reuses the victim's header value replaces the victim's cached query embedding before document scoring — so the victim's documents are scored against the **attacker's** query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim's `X-Request-Id`. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error.</p>
<p>This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models.</p>
<p>## Affected code</p>
<p>Links pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491c…</p></div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/ghsa-2phq-3phc-84px"/>
  </entry>
</feed>
