<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://vulnerability.circl.lu/rss/recent/all/10</id>
  <title>Most recent entries from all</title>
  <updated>2026-09-29T14:18:31.950412+00:00</updated>
  <author>
    <name>Vulnerability-Lookup</name>
    <email>info@circl.lu</email>
  </author>
  <link href="https://vulnerability.circl.lu" rel="alternate"/>
  <generator uri="https://lkiesow.github.io/python-feedgen" version="1.0.0">python-feedgen</generator>
  <subtitle>Contains only the most 10 recent entries.</subtitle>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/brew-acronym-cve-2026-81722</id>
    <title>BREW-acronym-CVE-2026-81722 — NLTK: Quadratic-time DoS in PorterStemmer via long runs of 'y'</title>
    <updated>2026-09-29T14:18:31.953056+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> Homebrew: acronym</p>
<p>`nltk.stem.PorterStemmer.stem()` -- a ubiquitous public API applied to arbitrary, often untrusted, tokens -- runs in O(n^2) time on a token containing a long run of the letter 'y', letting a single ~20-50 KB token pin a CPU core (CWE-407).</p>
<p>## Root cause</p>
<p>`_is_consonant(word, i)` was made *iterative* (commit for #3633, GHSA/CWE-674) to fix an earlier unbounded-recursion `RecursionError` on `'y'*10000`. The iterative form walks *backward* over the whole run of 'y's on every call:</p>
<p>```python
while i &gt; 0 and word[i] == 'y':
    negate = not negate
    i -= 1
```</p>
<p>`_measure()` then calls `_is_consonant(stem, i)` once for **every** position `i` of the stem. For a run of n 'y's that is sum_{i} O(i) = O(n^2). The recursion fix therefore traded a CWE-674 RecursionError for a CWE-407 quadratic-time DoS.</p>
<p>## Proof of concept</p>
<p>Measured (Python 3.13): `stem('y'*5000 + 'ness')` = 2.6s, `stem('y'*10000 + 'ness')` = 11.3s (2x input -&gt; ~4.3x time = quadratic), `stem('y'*20000 + 'ness')` &gt; 20s. A pure run of 'y' with no matching suffix is fast because the stemmer rules that call `_measure` do not fire; a real suffix such as 'ness' triggers `_measure` on the long stem.</p>
<p>```python
from nltk.stem import PorterStemmer
PorterStemmer().stem('y' * 20000 + 'ness')   # &gt;20s of CPU
```</p>
<p>## Impact</p>
<p>Stemming is routinely applied to untrusted text (search, indexing, NLP pipelines). A single unbroken ~20-50 KB token of 'y' characters (no whitespace, so it survives tokenization) causes multi-second-to-minu…</p></div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/brew-acronym-cve-2026-81722"/>
  </entry>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/cve-2026-81722</id>
    <title>CVE-2026-81722 — nltk PorterStemmer before 3.10.3 Quadratic-time DoS</title>
    <updated>2026-09-29T14:18:31.953151+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> nltk</p>
<p>nltk PorterStemmer in versions &lt;= 3.10.2 (fixed in 3.10.3) contains an inefficient-algorithmic-complexity denial of service in PorterStemmer.stem(). The _is_consonant() helper walks backward over the entire run of trailing 'y' characters on every call, and _measure() invokes it for each stem position, causing O(n^2) behavior. A single ~20-50 KB untrusted token consisting of a long run of the letter 'y' followed by a matching suffix (e.g., 'ness') can pin a CPU core for seconds to minutes, causing availability impact.</p></div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/cve-2026-81722"/>
  </entry>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/ghsa-ww6m-cw3f-q94g</id>
    <title>GHSA-ww6m-cw3f-q94g — NLTK: Quadratic-time DoS in PorterStemmer via long runs of 'y'</title>
    <updated>2026-09-29T14:18:31.953192+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> PyPI: nltk</p>
<p>`nltk.stem.PorterStemmer.stem()` -- a ubiquitous public API applied to arbitrary, often untrusted, tokens -- runs in O(n^2) time on a token containing a long run of the letter 'y', letting a single ~20-50 KB token pin a CPU core (CWE-407).</p>
<p>## Root cause</p>
<p>`_is_consonant(word, i)` was made *iterative* (commit for #3633, GHSA/CWE-674) to fix an earlier unbounded-recursion `RecursionError` on `'y'*10000`. The iterative form walks *backward* over the whole run of 'y's on every call:</p>
<p>```python
while i &gt; 0 and word[i] == 'y':
    negate = not negate
    i -= 1
```</p>
<p>`_measure()` then calls `_is_consonant(stem, i)` once for **every** position `i` of the stem. For a run of n 'y's that is sum_{i} O(i) = O(n^2). The recursion fix therefore traded a CWE-674 RecursionError for a CWE-407 quadratic-time DoS.</p>
<p>## Proof of concept</p>
<p>Measured (Python 3.13): `stem('y'*5000 + 'ness')` = 2.6s, `stem('y'*10000 + 'ness')` = 11.3s (2x input -&gt; ~4.3x time = quadratic), `stem('y'*20000 + 'ness')` &gt; 20s. A pure run of 'y' with no matching suffix is fast because the stemmer rules that call `_measure` do not fire; a real suffix such as 'ness' triggers `_measure` on the long stem.</p>
<p>```python
from nltk.stem import PorterStemmer
PorterStemmer().stem('y' * 20000 + 'ness')   # &gt;20s of CPU
```</p>
<p>## Impact</p>
<p>Stemming is routinely applied to untrusted text (search, indexing, NLP pipelines). A single unbroken ~20-50 KB token of 'y' characters (no whitespace, so it survives tokenization) causes multi-second-to-minu…</p></div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/ghsa-ww6m-cw3f-q94g"/>
  </entry>
</feed>
