<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0">
  <channel>
    <title>Most recent entries from all</title>
    <link>https://vulnerability.circl.lu</link>
    <description>Contains only the most 10 recent entries.</description>
    <docs>http://www.rssboard.org/rss-specification</docs>
    <generator>python-feedgen</generator>
    <language>en</language>
    <lastBuildDate>Thu, 01 Oct 2026 07:25:42 +0000</lastBuildDate>
    <item>
      <title>CVE-2026-73559 — vLLM: Completion prompt lists fan out into unbounded engine requests</title>
      <link>https://vulnerability.circl.lu/vuln/cve-2026-73559</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; vllm-project vllm&lt;/p&gt;
&lt;p&gt;vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[int]], prompt_to_seq() in vllm/renderers/inputs/preprocess.py and OnlineRenderer.preprocess_completion() in vllm/renderers/online_renderer.py expand every element, and vllm/entrypoints/openai/completion/serving.py creates one engine generator and response slot per prompt, allowing an authenticated API client to exhaust CPU, memory, async scheduling capacity, engine request slots, and response buffering with one request. This issue is fixed in version 0.26.0.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; vllm-project vllm&lt;/p&gt;
&lt;p&gt;vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[int]], prompt_to_seq() in vllm/renderers/inputs/preprocess.py and OnlineRenderer.preprocess_completion() in vllm/renderers/online_renderer.py expand every element, and vllm/entrypoints/openai/completion/serving.py creates one engine generator and response slot per prompt, allowing an authenticated API client to exhaust CPU, memory, async scheduling capacity, engine request slots, and response buffering with one request. This issue is fixed in version 0.26.0.&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://vulnerability.circl.lu/vuln/cve-2026-73559</guid>
    </item>
    <item>
      <title>PYSEC-2026-3704 — vLLM: Completion prompt lists fan out into unbounded engine requests</title>
      <link>https://vulnerability.circl.lu/vuln/pysec-2026-3704</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;## Summary&lt;/p&gt;
&lt;p&gt;The `/v1/completions` request model accepts `prompt` as a list of text prompts or a list of token-id prompts without any outer prompt-count bound. The serving path turns each element into a separate engine input, creates one engine generator per element, merges all generators, and allocates a response slot per prompt. An authenticated API client can therefore turn one request into an attacker-chosen number of backend subrequests before any aggregate request-count budget is enforced.&lt;/p&gt;
&lt;p&gt;## Technical Details&lt;/p&gt;
&lt;p&gt;`CompletionRequest.prompt` allows both list-shaped prompt inputs and scalar prompts:&lt;/p&gt;
&lt;p&gt;```python
# vllm/entrypoints/openai/completion/protocol.py
prompt: (
    list[Annotated[int, Field(ge=0)]]
    | list[list[Annotated[int, Field(ge=0)]]]
    | str
    | list[str]
    | None
) = None
```&lt;/p&gt;
&lt;p&gt;The validator only requires some prompt-like input to be present:&lt;/p&gt;
&lt;p&gt;```python
def validate_prompt_and_prompt_embeds(cls, data):
    prompt = data.get(&amp;#34;prompt&amp;#34;)
    prompt_embeds = data.get(&amp;#34;prompt_embeds&amp;#34;)
    ...
    if prompt_is_empty and embeds_is_empty:
        raise VLLMValidationError(...)
```&lt;/p&gt;
&lt;p&gt;The renderer then expands list-shaped prompts as a sequence. `prompt_to_seq()` wraps a scalar string or a single token-id list, but returns a `list[str]` or `list[list[int]]` unchanged:&lt;/p&gt;
&lt;p&gt;```python
# vllm/renderers/inputs/preprocess.py
def prompt_to_seq(prompt_or_prompts):
    if isinstance(prompt_or_prompts, (dict, str, bytes)) or (
        len(prompt_or_prompts) &amp;gt; 0 and is_list_of(…&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: vllm&lt;/p&gt;
&lt;p&gt;## Summary&lt;/p&gt;
&lt;p&gt;The `/v1/completions` request model accepts `prompt` as a list of text prompts or a list of token-id prompts without any outer prompt-count bound. The serving path turns each element into a separate engine input, creates one engine generator per element, merges all generators, and allocates a response slot per prompt. An authenticated API client can therefore turn one request into an attacker-chosen number of backend subrequests before any aggregate request-count budget is enforced.&lt;/p&gt;
&lt;p&gt;## Technical Details&lt;/p&gt;
&lt;p&gt;`CompletionRequest.prompt` allows both list-shaped prompt inputs and scalar prompts:&lt;/p&gt;
&lt;p&gt;```python
# vllm/entrypoints/openai/completion/protocol.py
prompt: (
    list[Annotated[int, Field(ge=0)]]
    | list[list[Annotated[int, Field(ge=0)]]]
    | str
    | list[str]
    | None
) = None
```&lt;/p&gt;
&lt;p&gt;The validator only requires some prompt-like input to be present:&lt;/p&gt;
&lt;p&gt;```python
def validate_prompt_and_prompt_embeds(cls, data):
    prompt = data.get(&amp;#34;prompt&amp;#34;)
    prompt_embeds = data.get(&amp;#34;prompt_embeds&amp;#34;)
    ...
    if prompt_is_empty and embeds_is_empty:
        raise VLLMValidationError(...)
```&lt;/p&gt;
&lt;p&gt;The renderer then expands list-shaped prompts as a sequence. `prompt_to_seq()` wraps a scalar string or a single token-id list, but returns a `list[str]` or `list[list[int]]` unchanged:&lt;/p&gt;
&lt;p&gt;```python
# vllm/renderers/inputs/preprocess.py
def prompt_to_seq(prompt_or_prompts):
    if isinstance(prompt_or_prompts, (dict, str, bytes)) or (
        len(prompt_or_prompts) &amp;gt; 0 and is_list_of(…&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://vulnerability.circl.lu/vuln/pysec-2026-3704</guid>
    </item>
  </channel>
</rss>
